Fetching the paper…
Reading the bibliography…
Recent techniques for approximating Nash equilibria in very large games leverage neural networks to learn approximately optimal policies (strategies).
Low-variance and zero-variance baselines for extensive-form games
Davis, T., Schmid, M., and Bowling, M · 1907
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
A Course in Game Theory
Osborne, M. J. and Rubinstein, A · 1994
Earlier work this paper cites.
On-line Q-learning using connectionist systems , volume 37
Rummery, G. A. and Niranjan, M · 1994
Earlier work this paper cites.
A simple adaptive procedure leading to correlated equilibrium
Hart, S. and Mas-Colell, A · 2000
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Brafman, R. I. and Tennenholtz, M · 2002
Earlier work this paper cites.
Dynamic programming for partially observable stochastic games
Hansen, E. A., Bernstein, D. S., and Zilberstein, S · 2004
Earlier work this paper cites.
Monte carlo sampling for regret minimization in extensive games
Lanctot, M., Waugh, K., Zinkevich, M., and Bowling, M · 2009
Earlier work this paper cites.
Efficient monte carlo counterfactual regret minimization in games with many player actions
Burch, N., Lanctot, M., Szafron, D., and Gibson, R · 2012
Earlier work this paper cites.
Generalized sampling and variance in counterfactual regret minimization
Gibson, R., Lanctot, M., Burch, N., Szafron, D., and Bowling, M · 2012
Earlier work this paper cites.
Solving imperfect information games using decomposition
Burch, N., Johanson, M., and Bowling, M · 2014
Earlier work this paper cites.
Heads-up limit hold’em poker is solved
Bowling, M., Burch, N., Johanson, M., and Tammelin, O · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Deep reinforcement learning from self-play in imperfect-information games
Heinrich, J. and Silver, D · 2016
Earlier work this paper cites.
Refining subgames in large imperfect information games
Moravcik, M., Schmid, M., Ha, K., Hladik, M., and Gaukrodger, S · 2016
Earlier work this paper cites.
A unified game-theoretic approach to multiagent reinforcement learning
Lanctot, M., Zambaldi, V., Gruslys, A., Lazaridou, A., Tuyls, K., Pérolat, J., Silver, D., and Graepel, T · 2017
Earlier work this paper cites.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Moravčík, M., Schmid, M., Burch, N., Lisỳ, V., Morrill, D., Bard, N., Davis, T., Waugh, K., Johanson, M., and Bowling, M · 2017
Earlier work this paper cites.
Robust adversarial reinforcement learning
Pinto, L., Davidson, J., Sukthankar, R., and Gupta, A · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Earlier work this paper cites.
Online reinforcement learning in stochastic games
Wei, C.-Y., Hong, Y.-T., and Lu, C.-J · 2017
Earlier work this paper cites.
Superhuman AI for heads-up no-limit poker: Libratus beats top professionals
Brown, N. and Sandholm, T · 2018
Earlier work this paper cites.
Depth-limited solving for imperfect-information games
Brown, N., Sandholm, T., and Amos, B · 2018
Cited alongside, same era.
Rllib: Abstractions for distributed reinforcement learning
Liang, E., Liaw, R., Nishihara, R., Moritz, P., Fox, R., Goldberg, K., Gonzalez, J., Jordan, M., and Stoica, I · 2018
Cited alongside, same era.
Ray: A distributed framework for emerging { \{ AI } \} applications
Moritz, P., Nishihara, R., Wang, S., Tumanov, A., Liaw, R., Liang, E., Elibol, M., Yang, Z., Paul, W., Jordan, M. I., et al · 2018
Cited alongside, same era.
Actor-critic fictitious play in simultaneous move multistage games
Perolat, J., Piot, B., and Pietquin, O · 2018
Cited alongside, same era.
Actor-critic policy optimization in partially observable multiagent environments
Srinivasan, S., Lanctot, M., Zambaldi, V., Pérolat, J., Tuyls, K., Munos, R., and Bowling, M · 2018
Cited alongside, same era.
Superhuman AI for multiplayer poker
Pipeline PSRO: A scalable approach for finding approximate Nash equilibria in large games
McAleer, S., Lanier, J., Fox, R., and Baldi, P · 2020
Later among the works it cites.
DREAM: Deep regret minimization with advantage baselines and model-free learning
Steinberger, E., Lerer, A., and Brown, N · 2020
Later among the works it cites.
Learning zero-sum simultaneous-move markov games using function approximation and correlated equilibrium
Xie, Q., Chen, Y., Wang, Z., and Yang, Z · 2020
Later among the works it cites.
Discovering multi-agent auto-curricula in two-player zero-sum games
Feng, X., Slumbers, O., Yang, Y., Wan, Z., Liu, B., McAleer, S., Wen, Y., and Wang, J · 2021
Later among the works it cites.
V-learning–a simple, efficient, decentralized algorithm for multiagent rl
Jin, C., Liu, Q., Wang, Y., and Yu, T · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Brown, N. and Sandholm, T · 2019
Cited alongside, same era.
Deep counterfactual regret minimization
Brown, N., Lerer, A., Gross, S., and Sandholm, T · 2019
Cited alongside, same era.
Openspiel: A framework for reinforcement learning in games
Lanctot, M., Lockhart, E., Lespiau, J.-B., Zambaldi, V., Upadhyay, S., Pérolat, J., Srinivasan, S., Timbers, F., Tuyls, K., Omidshafiei, S., et al · 2019
Cited alongside, same era.
Double neural counterfactual regret minimization
Li, H., Hu, K., Zhang, S., Qi, Y., and Song, L · 2019
Cited alongside, same era.
A generalized training approach for multiagent learning
Muller, P., Omidshafiei, S., Rowland, M., Tuyls, K., Perolat, J., Liu, S., Hennes, D., Marris, L., Lanctot, M., Hughes, E., et al · 2019
Cited alongside, same era.
Variance reduction in monte carlo counterfactual regret minimization (VR-MCCFR) for extensive form games using baselines
Schmid, M., Burch, N., Lanctot, M., Moravcik, M., Kadlec, R., and Bowling, M · 2019
Cited alongside, same era.
Finding friend and foe in multi-agent games
Serrino, J., Kleiman-Weiner, M., Parkes, D. C., and Tenenbaum, J · 2019
Cited alongside, same era.
Global convergence of multi-agent policy gradient in markov potential games
Leonardos, S., Overman, W., Panageas, I., and Piliouras, G · 2021
Later among the works it cites.
XDO: A double oracle algorithm for extensive-form games
McAleer, S., Lanier, J., Baldi, P., and Fox, R · 2021
Later among the works it cites.
Learning in nonzero-sum stochastic games with potentials
Mguni, D. H., Wu, Y., Du, Y., Yang, Y., Wang, Z., Li, M., Wen, Y., Jennings, J., and Wang, J · 2021
Later among the works it cites.
From Poincaré recurrence to convergence in imperfect information games: Finding equilibrium via regularization
Perolat, J., Munos, R., Lespiau, J.-B., Omidshafiei, S., Rowland, M., Ortega, P., Burch, N., Anthony, T., Balduzzi, D., De Vylder, B., et al · 2021
Later among the works it cites.
Schmid, M., Moravcik, M., Burch, N., Kadlec, R., Davidson, J., Waugh, K., Bard, N., Timbers, F., Lanctot, M., Holland, Z., et al · 2021
Later among the works it cites.
Douzero: Mastering doudizhu with self-play deep reinforcement learning
Zha, D., Xie, J., Ma, W., Zhang, S., Lian, X., Hu, X., and Liu, J · 2021
Later among the works it cites.
Subgame solving without common knowledge
Zhang, B. and Sandholm, T · 2021
Later among the works it cites.
Gradient play in stochastic games: stationary points, convergence, and sample complexity
Zhang, R., Ren, Z., and Li, N · 2021
Later among the works it cites.
Ding, D., Wei, C.-Y., Zhang, K., and Jovanović, M. R · 2022
Closest in time.
Independent natural policy gradient always converges in markov potential games
Fox, R., Mcaleer, S. M., Overman, W., and Panageas, I · 2022
Closest in time.
Actor-critic policy optimization in a large-scale imperfect-information game
Fu, H., Liu, W., Wu, S., Wang, Y., Yang, T., Li, K., Xing, J., Li, B., Ma, B., Fu, Q., and Wei, Y · 2022
Closest in time.
Rethinking formal models of partially observable multiagent decision making
Kovařík, V., Schmid, M., Burch, N., Bowling, M., and Lisỳ, V · 2022
Closest in time.
Model-free neural counterfactual regret minimization with bootstrap learning
Liu, W., Li, B., and Togelius, J · 2022
Closest in time.
Anytime optimal psro for two-player zero-sum games
McAleer, S., Wang, K., Lanctot, M., Lanier, J., Baldi, P., and Fox, R · 2022
Closest in time.
Mastering the game of stratego with model-free multiagent reinforcement learning
Perolat, J., de Vylder, B., Hennes, D., Tarassov, E., Strub, F., de Boer, V., Muller, P., Connor, J. T., Burch, N., Anthony, T., et al · 2022
Closest in time.
Outracing champion gran turismo drivers with deep reinforcement learning
Wurman, P. R., Barrett, S., Kawamoto, K., MacGlashan, J., Subramanian, K., Walsh, T. J., Capobianco, R., Devlic, A., Eckert, F., Fuchs, F., et al · 2022
Closest in time.