Fetching the paper…
Reading the bibliography…
Policy space response oracles (PSRO) is a multi-agent reinforcement learning algorithm that has achieved state-of-the-art performance in very large two-player zero-sum games.
Gambling in a rigged casino: The adversarial multi-armed bandit problem
Auer, P., Cesa-Bianchi, N., Freund, Y., and Schapire, R. E · 1995
Earlier work this paper cites.
Regret minimization in games with incomplete information
Zinkevich, M., Johanson, M., Bowling, M., and Piccione, C · 1995
Earlier work this paper cites.
Adaptive game playing using multiplicative weights
Freund, Y. and Schapire, R. E · 1999
Earlier work this paper cites.
A simple adaptive procedure leading to correlated equilibrium
Hart, S. and Mas-Colell, A · 2000
Earlier work this paper cites.
Planning in the presence of cost functions controlled by an adversary
McMahan, H. B., Gordon, G. J., and Blum, A · 2003
Earlier work this paper cites.
Dynamic programming for partially observable stochastic games
Hansen, E. A., Bernstein, D. S., and Zilberstein, S · 2004
Earlier work this paper cites.
Prediction, learning, and games
Cesa-Bianchi, N. and Lugosi, G · 2006
Earlier work this paper cites.
Methods for empirical game-theoretic analysis
Wellman, M. P · 2006
Earlier work this paper cites.
A new algorithm for generating equilibria in massive zero-sum games
Zinkevich, M., Bowling, M., and Burch, N · 2007
Earlier work this paper cites.
On range of skill
Hansen, T. D., Miltersen, P. B., and Sørensen, T. B · 2008
Earlier work this paper cites.
Strategy exploration in empirical games
Jordan, P. R., Schvartzman, L. J., and Wellman, M. P · 2010
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Bubeck, S. and Cesa-Bianchi, N · 2012
Cited alongside, same era.
Finding optimal abstract strategies in extensive form games
Johanson, M., Bard, N., Burch, N., and Bowling, M · 2012
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D · 2016
Cited alongside, same era.
A unified game-theoretic approach to multiagent reinforcement learning
Lanctot, M., Zambaldi, V., Gruslys, A., Lazaridou, A., Tuyls, K., Pérolat, J., Silver, D., and Graepel, T · 2017
Cited alongside, same era.
Stable-predictive optimistic counterfactual regret minimization
Farina, G., Kroer, C., Brown, N., and Sandholm, T · 2019
Later among the works it cites.
Openspiel: A framework for reinforcement learning in games
Lanctot, M., Lockhart, E., Lespiau, J.-B., Zambaldi, V., Upadhyay, S., Pérolat, J., Srinivasan, S., Timbers, F., Tuyls, K., Omidshafiei, S., et al · 2019
Later among the works it cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Later among the works it cites.
Iterated deep reinforcement learning in games: History-aware training for improved stability
Wright, M., Wang, Y., and Wellman, M. P · 2019
Later among the works it cites.
Pipeline PSRO: A scalable approach for finding approximate Nash equilibria in large games
McAleer, S., Lanier, J., Fox, R., and Baldi, P · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Rllib: Abstractions for distributed reinforcement learning
Liang, E., Liaw, R., Nishihara, R., Moritz, P., Fox, R., Goldberg, K., Gonzalez, J., Jordan, M., and Stoica, I · 2018
Cited alongside, same era.
Open-ended learning in symmetric zero-sum games
Balduzzi, D., Garnelo, M., Bachrach, Y., Czarnecki, W., Perolat, J., Jaderberg, M., and Graepel, T · 2019
Cited alongside, same era.
Solving imperfect-information games via discounted regret minimization
Brown, N. and Sandholm, T · 2019
Cited alongside, same era.
Discovering multi-agent auto-curricula in two-player zero-sum games
Feng, X., Slumbers, O., Yang, Y., Wan, Z., Liu, B., McAleer, S., Wen, Y., and Wang, J · 2021
Later among the works it cites.
XDO: A double oracle algorithm for extensive-form games
McAleer, S., Lanier, J., Baldi, P., and Fox, R · 2021
Later among the works it cites.
Iterative empirical game solving via single policy best response
Smith, M. O., Anthony, T., and Wellman, M. P · 2021
Later among the works it cites.
Evaluating strategy exploration in empirical game-theoretic analysis
Wang, Y., Ma, Q., and Wellman, M. P · 2021
Later among the works it cites.