Sample-Efficient Policy Space Response Oracles with Joint Experience Best Response
Files
Publication date
2026-05-24
Authors
Bighashdel, Ariyan
Simão, Thiago D.
Oliehoek, Frans A.
Editors
Advisors
Supervisors
Document Type
Part of book
Metadata
Show full item recordCollections
License
cc_by
Abstract
Multi-agent reinforcement learning (MARL) offers a scalable alternative to exact game-theoretic analysis but suffers from non-stationarity and the need to maintain diverse populations of strategies that capture non-transitive interactions. Policy Space Response Oracles (PSRO) address these issues by iteratively expanding a restricted game with approximate best responses (BRs), yet per-agent BR training makes it prohibitively expensive in many-agent or simulator-expensive settings. We introduce Joint Experience Best Response (JBR), a drop-in modification to PSRO that collects trajectories once under the current meta-strategy profile and reuses this joint dataset to compute BRs for all agents simultaneously. This amortizes environment interaction and improves the sample efficiency of best-response computation. Because JBR converts BR computation into an offline RL problem, we propose three remedies for distribution-shift bias: (i) Conservative JBR with safe policy improvement, (ii) Exploration-Augmented JBR that perturbs data collection and admits theoretical guarantees, and (iii) Hybrid BR that interleaves JBR with periodic independent BR updates. Across benchmark multi-agent environments, Exploration-Augmented JBR achieves the best accuracy-efficiency trade-off, while Hybrid BR attains near-PSRO performance at a fraction of the sample cost. Overall, JBR makes PSRO substantially more practical for large-scale strategic learning while preserving equilibrium robustness.
Keywords
Game Theory, Multi-Agent Reinforcement Learning, Policy Space Response Oracles, Artificial Intelligence
Citation
Bighashdel, A, Simão, T D & Oliehoek, F A 2026, Sample-Efficient Policy Space Response Oracles with Joint Experience Best Response. in AAMAS 2026 - Proceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems. AAMAS 2026 - Proceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems, Association for Computing Machinery, pp. 2401-2409, 25th International Conference on Autonomous Agents and Multiagent Systems, AAMAS 2026, Paphos, Cyprus, 25/05/26. https://doi.org/10.65109/BQIF3470, conference