Fetching the paper…
Reading the bibliography…
Finding approximate Nash equilibria in zero-sum imperfect-information games is challenging when the number of information states is large.
Evolutionary stable strategies and game dynamics
P. D. Taylor and L. B. Jonker · 1978
Earlier work this paper cites.
The theory of learning in games
D. Fudenberg, F. Drew, D. K. Levine, and D. K. Levine · 1998
Earlier work this paper cites.
Multiagent learning using a variable learning rate
M. Bowling and M. Veloso · 2002
Earlier work this paper cites.
Planning in the presence of cost functions controlled by an adversary
H. B. McMahan, G. J. Gordon, and A. Blum · 2003
Earlier work this paper cites.
Algorithmic Game Theory
N. Nisan, T. Roughgarden, E. Tardos, and V. V. Vazirani · 2007
Earlier work this paper cites.
Multiagent systems: Algorithmic, game-theoretic, and logical foundations
Y. Shoham and K. Leyton-Brown · 2008
Earlier work this paper cites.
Regret minimization in games with incomplete information
M. Zinkevich, M. Johanson, M. Bowling, and C. Piccione · 2008
Earlier work this paper cites.
The 3rd stratego computer world championship
S. Jug and M. Schadd · 2009
Earlier work this paper cites.
Quiescence search for stratego
M. Schadd and M. Winands · 2009
Earlier work this paper cites.
Fictitious self-play in extensive-form games
J. Heinrich, M. Lanctot, and D. Silver · 2015
Earlier work this paper cites.
Deep reinforcement learning from self-play in imperfect-information games
J. Heinrich and D. Silver · 2016
Cited alongside, same era.
A unified game-theoretic approach to multiagent reinforcement learning
M. Lanctot, V. Zambaldi, A. Gruslys, A. Lazaridou, K. Tuyls, J. Pérolat, D. Silver, and T. Graepel · 2017
Cited alongside, same era.
Deep reinforcement learning: An overview
Y. Li · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, Y. Chen, T. Lillicrap, F. Hui, L. Sifre, G. v. d. Driessche, T. Graepel, and D. Hassabis · 2017
Cited alongside, same era.
Superhuman ai for heads-up no-limit poker: Libratus beats top professionals
N. Brown and T. Sandholm · 2018
Cited alongside, same era.
Open-ended learning in symmetric zero-sum games
D. Balduzzi, M. Garnelo, Y. Bachrach, W. Czarnecki, J. Perolat, M. Jaderberg, and T. Graepel · 2019
Later among the works it cites.
Dota 2 with large scale deep reinforcement learning
C. Berner, G. Brockman, B. Chan, V. Cheung, P. Dębiak, C. Dennison, D. Farhi, Q. Fischer, S. Hashme, C. Hesse, et al · 2019
Later among the works it cites.
Deep counterfactual regret minimization
N. Brown, A. Lerer, S. Gross, and T. Sandholm · 2019
Later among the works it cites.
Soft actor-critic for discrete action settings
P. Christodoulou · 2019
Later among the works it cites.
Human-level performance in 3d multiplayer games with population-based reinforcement learning
M. Jaderberg, W. M. Czarnecki, I. Dunning, L. Marris, G. Lever, A. G. Castaneda, C. Beattie, N. C. Rabinowitz, A. S. Morcos, A. Ruderman, et al · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
Rllib: Abstractions for distributed reinforcement learning
E. Liang, R. Liaw, R. Nishihara, P. Moritz, R. Fox, K. Goldberg, J. Gonzalez, M. Jordan, and I. Stoica · 2018
Cited alongside, same era.
Ray: A distributed framework for emerging { \{ AI } \} applications
P. Moritz, R. Nishihara, S. Wang, A. Tumanov, R. Liaw, E. Liang, M. Elibol, Z. Yang, W. Paul, M. I. Jordan, et al · 2018
Cited alongside, same era.
Actor-critic policy optimization in partially observable multiagent environments
S. Srinivasan, M. Lanctot, V. Zambaldi, J. Pérolat, K. Tuyls, R. Munos, and M. Bowling · 2018
Cited alongside, same era.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, et al · 2019
Later among the works it cites.
The advantage regret-matching actor-critic
A. Gruslys, M. Lanctot, R. Munos, F. Timbers, M. Schmid, J. Perolat, D. Morrill, V. Zambaldi, J.-B. Lespiau, J. Schultz, et al · 2020
Closest in time.
A generalized training approach for multiagent learning
P. Muller, S. Omidshafiei, M. Rowland, K. Tuyls, J. Perolat, S. Liu, D. Hennes, L. Marris, M. Lanctot, E. Hughes, et al · 2020
Closest in time.
Dream: Deep regret minimization with advantage baselines and model-free learning
E. Steinberger, A. Lerer, and N. Brown · 2020
Closest in time.