Fetching the paper…
Reading the bibliography…
In competitive two-agent environments, deep reinforcement learning (RL) methods based on the \emph{Double Oracle (DO)} algorithm, such as \emph{Policy Space Response Oracles (PSRO)} and \emph{Anytime PSRO (APSRO)}, iteratively add RL best response policies to a population.
Iterative solution of games by fictitious play
Brown, G. W · 1951
Earlier work this paper cites.
Random sampling with a reservoir
Vitter, J. S · 1985
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Brafman, R. I. and Tennenholtz, M · 2002
Earlier work this paper cites.
Planning in the presence of cost functions controlled by an adversary
McMahan, H. B., Gordon, G. J., and Blum, A · 2003
Earlier work this paper cites.
Dynamic programming for partially observable stochastic games
Hansen, E. A., Bernstein, D. S., and Zilberstein, S · 2004
Earlier work this paper cites.
Methods for empirical game-theoretic analysis
Wellman, M. P · 2006
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Deep reinforcement learning from self-play in imperfect-information games
Heinrich, J. and Silver, D · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D · 2016
Earlier work this paper cites.
A unified game-theoretic approach to multiagent reinforcement learning
Lanctot, M., Zambaldi, V., Gruslys, A., Lazaridou, A., Tuyls, K., Pérolat, J., Silver, D., and Graepel, T · 2017
Earlier work this paper cites.
Online reinforcement learning in stochastic games
Wei, C.-Y., Hong, Y.-T., and Lu, C.-J · 2017
Earlier work this paper cites.
Rllib: Abstractions for distributed reinforcement learning
Liang, E., Liaw, R., Nishihara, R., Moritz, P., Fox, R., Goldberg, K., Gonzalez, J., Jordan, M., and Stoica, I · 2018
Earlier work this paper cites.
Actor-critic fictitious play in simultaneous move multistage games
Perolat, J., Piot, B., and Pietquin, O · 2018
Earlier work this paper cites.
Actor-critic policy optimization in partially observable multiagent environments
Srinivasan, S., Lanctot, M., Zambaldi, V., Pérolat, J., Tuyls, K., Munos, R., and Bowling, M · 2018
Cited alongside, same era.
Deep counterfactual regret minimization
Brown, N., Lerer, A., Gross, S., and Sandholm, T · 2019
Cited alongside, same era.
Openspiel: A framework for reinforcement learning in games
Lanctot, M., Lockhart, E., Lespiau, J.-B., Zambaldi, V., Upadhyay, S., Pérolat, J., Srinivasan, S., Timbers, F., Tuyls, K., Omidshafiei, S., et al · 2019
Cited alongside, same era.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Cited alongside, same era.
Independent policy gradient methods for competitive reinforcement learning
Daskalakis, C., Foster, D. J., and Golowich, N · 2020
Cited alongside, same era.
Global convergence of multi-agent policy gradient in markov potential games
Leonardos, S., Overman, W., Panageas, I., and Piliouras, G · 2021
Later among the works it cites.
Towards unifying behavioral and response diversity for open-ended learning in zero-sum games
Liu, X., Jia, H., Wen, Y., Hu, Y., Chen, Y., Fan, C., Hu, Z., and Yang, Y · 2021
Later among the works it cites.
Multi-agent training beyond zero-sum with correlated equilibrium meta-solvers
Marris, L., Muller, P., Lanctot, M., Tuyls, K., and Graepel, T · 2021
Later among the works it cites.
XDO: A double oracle algorithm for extensive-form games
McAleer, S., Lanier, J. B., Wang, K. A., Baldi, P., and Fox, R · 2021
Later among the works it cites.
Learning in nonzero-sum stochastic games with potentials
Mguni, D. H., Wu, Y., Du, Y., Yang, Y., Wang, Z., Li, M., Wen, Y., Jennings, J., and Wang, J · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neural replicator dynamics: Multiagent learning via hedging policy gradients
Hennes, D., Morrill, D., Omidshafiei, S., Munos, R., Perolat, J., Lanctot, M., Gruslys, A., Lespiau, J.-B., Parmas, P., Duéñez-Guzmán, E., et al · 2020
Cited alongside, same era.
Pipeline PSRO: A scalable approach for finding approximate Nash equilibria in large games
McAleer, S., Lanier, J., Fox, R., and Baldi, P · 2020
Cited alongside, same era.
A generalized training approach for multiagent learning
Muller, P., Omidshafiei, S., Rowland, M., Tuyls, K., Perolat, J., Liu, S., Hennes, D., Marris, L., Lanctot, M., Hughes, E., et al · 2020
Cited alongside, same era.
Dream: Deep regret minimization with advantage baselines and model-free learning
Steinberger, E., Lerer, A., and Brown, N · 2020
Cited alongside, same era.
Learning zero-sum simultaneous-move markov games using function approximation and correlated equilibrium
Xie, Q., Chen, Y., Wang, Z., and Yang, Z · 2020
Cited alongside, same era.
Dinh, L. C., Yang, Y., Tian, Z., Nieves, N. P., Slumbers, O., Mguni, D. H., Ammar, H. B., and Wang, J · 2021
Cited alongside, same era.
Neural auto-curricula in two-player zero-sum games
Feng, X., Slumbers, O., Wan, Z., Liu, B., McAleer, S., Wen, Y., Wang, J., and Yang, Y · 2021
Cited alongside, same era.
Modelling behavioural diversity for learning in open-ended games
Perez-Nieves, N., Yang, Y., Slumbers, O., Mguni, D. H., Wen, Y., and Wang, J · 2021
Later among the works it cites.
From Poincaré recurrence to convergence in imperfect information games: Finding equilibrium via regularization
Perolat, J., Munos, R., Lespiau, J.-B., Omidshafiei, S., Rowland, M., Ortega, P., Burch, N., Anthony, T., Balduzzi, D., De Vylder, B., et al · 2021
Later among the works it cites.
Gradient play in stochastic games: stationary points, convergence, and sample complexity
Zhang, R., Ren, Z., and Li, N · 2021
Later among the works it cites.
Ding, D., Wei, C.-Y., Zhang, K., and Jovanović, M. R · 2022
Closest in time.
Independent natural policy gradient always converges in markov potential games
Fox, R., Mcaleer, S. M., Overman, W., and Panageas, I · 2022
Closest in time.
Mastering the game of stratego with model-free multiagent reinforcement learning
Perolat, J., de Vylder, B., Hennes, D., Tarassov, E., Strub, F., de Boer, V., Muller, P., Connor, J. T., Burch, N., Anthony, T., et al · 2022
Closest in time.
Learning risk-averse equilibria in multi-agent systems
Slumbers, O., Mguni, D. H., McAleer, S., Wang, J., and Yang, Y · 2022
Closest in time.
Sokota, S., D’Orazio, R., Kolter, J. Z., Loizou, N., Lanctot, M., Mitliagkas, I., Brown, N., and Kroer, C · 2022
Closest in time.