Fetching the paper…
Reading the bibliography…
Extensive-form games provide a versatile framework for modeling interactions of multiple agents subjected to imperfect observations and stochastic events.
Iterative solution of games by fictitious play
Brown, G. W · 1951
Earlier work this paper cites.
Planning in the presence of cost functions controlled by an adversary
McMahan, H. B., Gordon, G. J., and Blum, A · 2003
Earlier work this paper cites.
Bayes’ bluff: opponent modelling in poker
Southey, F., Bowling, M., Larson, B., Piccione, C., Burch, N., Billings, D., and Rayner, C · 2005
Earlier work this paper cites.
Regret minimization in games with incomplete information
Zinkevich, M., Johanson, M., Bowling, M., and Piccione, C · 2007
Earlier work this paper cites.
Multiagent Systems: Algorithmic, Game-Theoretic, and Logical Foundations
Shoham, Y. and Leyton-Brown, K · 2008
Earlier work this paper cites.
Abstraction pathologies in extensive games
Waugh, K., Schnizlein, D., Bowling, M. H., and Szafron, D · 2009
Earlier work this paper cites.
A double oracle algorithm for zero-sum security games on graphs
Jain, M., Korzhyk, D., Vaněk, O., Conitzer, V., Pěchouček, M., and Tambe, M · 2011
Earlier work this paper cites.
Security and Game Theory: Algorithms, Deployed Systems, Lessons Learned
Tambe, M · 2011
Earlier work this paper cites.
Generalized sampling and variance in counterfactual regret minimization
Gibson, R., Lanctot, M., Burch, N., Szafron, D., and Bowling, M · 2012
Earlier work this paper cites.
Heads-up limit Hold’em poker is solved
Bowling, M., Burch, N., Johanson, M., and Tammelin, O · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Fictitious self-play in extensive-form games
Heinrich, J., Lanctot, M., and Silver, D · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Deep reinforcement learning from self-play in imperfect-information games
Heinrich, J. and Silver, D · 2016
Cited alongside, same era.
A unified game-theoretic approach to multiagent reinforcement learning
Lanctot, M., Zambaldi, V., Gruslys, A., Lazaridou, A., Tuyls, K., Pérolat, J., Silver, D., and Graepel, T · 2017
Cited alongside, same era.
DeepStack: Expert-level artificial intelligence in heads-up no-limit poker
Moravčík, M., Schmid, M., Burch, N., Lisỳ, V., Morrill, D., Bard, N., Davis, T., Waugh, K., Johanson, M., and Bowling, M · 2017
Cited alongside, same era.
Human-level performance in no-press diplomacy via equilibrium search
Gray, J., Lerer, A., Bakhtin, A., and Brown, N · 2020
Later among the works it cites.
Pipeline PSRO: A scalable approach for finding approximate nash equilibria in large games
Mcaleer, S., Lanier, J., Fox, R., and Baldi, P · 2020
Later among the works it cites.
A generalized training approach for multiagent learning
Muller, P., Omidshafiei, S., Rowland, M., Tuyls, K., Perolat, J., Liu, S., Hennes, D., Marris, L., Lanctot, M., Hughes, E., et al · 2020
Later among the works it cites.
Iterative empirical game solving via single policy best response
Smith, M., Anthony, T., and Wellman, M · 2020
Later among the works it cites.
DREAM: Deep regret minimization with advantage baselines and model-free learning
Steinberger, E., Lerer, A., and Brown, N · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Superhuman ai for heads-up no-limit poker: Libratus beats top professionals
Brown, N. and Sandholm, T · 2018
Cited alongside, same era.
Actor-critic policy optimization in partially observable multiagent environments
Srinivasan, S., Lanctot, M., Zambaldi, V., Pérolat, J., Tuyls, K., Munos, R., and Bowling, M · 2018
Cited alongside, same era.
Deep counterfactual regret minimization
Brown, N., Lerer, A., Gross, S., and Sandholm, T · 2019
Cited alongside, same era.
OpenSpiel: A framework for reinforcement learning in games
Lanctot, M., Lockhart, E., Lespiau, J.-B., Zambaldi, V., Upadhyay, S., Pérolat, J., Srinivasan, S., Timbers, F., Tuyls, K., Omidshafiei, S., et al · 2019
Cited alongside, same era.
Single deep counterfactual regret minimization
Steinberger, E · 2019
Cited alongside, same era.
Fictitious play outperforms counterfactual regret minimization
Ganzfried, S · 2020
Cited alongside, same era.
Solving imperfect-information games via discounted regret minimization
Brown, N. and Sandholm, T
Cited in the paper.
Bakhtin, A., Wu, D., Lerer, A., and Brown, N · 2021
Later among the works it cites.
Discovering multi-agent auto-curricula in two-player zero-sum games
Feng, X., Slumbers, O., Yang, Y., Wan, Z., Liu, B., McAleer, S., Wen, Y., and Wang, J · 2021
Later among the works it cites.
Multi-agent training beyond zero-sum with correlated equilibrium meta-solvers
Marris, L., Muller, P., Lanctot, M., Tuyls, K., and Grapael, T · 2021
Later among the works it cites.
Learning equilibria in mean-field games: Introducing mean-field PSRO
Muller, P., Rowland, M., Elie, R., Piliouras, G., Perolat, J., Lauriere, M., Marinier, R., Pietquin, O., and Tuyls, K · 2021
Later among the works it cites.
Modelling behavioural diversity for learning in open-ended games
Perez-Nieves, N., Yang, Y., Slumbers, O., Mguni, D. H., Wen, Y., and Wang, J · 2021
Later among the works it cites.
Schmid, M., Moravcik, M., Burch, N., Kadlec, R., Davidson, J., Waugh, K., Bard, N., Timbers, F., Lanctot, M., Holland, Z., et al · 2021
Later among the works it cites.