Fetching the paper…
Reading the bibliography…
Policy Space Response Oracles (PSRO) is a reinforcement learning (RL) algorithm for two-player zero-sum games that has been empirically shown to find approximate Nash equilibria in large games.
OpenSpiel: A framework for reinforcement learning in games
Lanctot, M., Lockhart, E., Lespiau, J.-B., Zambaldi, V., Upadhyay, S., Pérolat, J., Srinivasan, S., Timbers, F., Tuyls, K., Omidshafiei, S., Hennes, D., Morrill, D., Muller, P., Ewalds, T., Faulkner, R., Kramár, J., Vylder, B. D., Saeta, B., Bradbury, J., Ding, D., Borgeaud, S., Lai, M., Schrittwieser, J., Anthony, T., Hughes, E., Danihelka, I., and Ryan-Davis, J · 1908
Earlier work this paper cites.
Openspiel: A framework for reinforcement learning in games
Lanctot, M., Lockhart, E., Lespiau, J.-B., Zambaldi, V., Upadhyay, S., Pérolat, J., Srinivasan, S., Timbers, F., Tuyls, K., Omidshafiei, S., et al · 1908
Earlier work this paper cites.
Iterative solution of games by fictitious play
Brown, G. W · 1951
Earlier work this paper cites.
Contributions to the Theory of Games , volume 2
Kuhn, H. W. and Tucker, A. W · 1953
Earlier work this paper cites.
Regret-based pruning in extensive-form games
Brown, N. and Sandholm, T · 1980
Earlier work this paper cites.
A simple adaptive procedure leading to correlated equilibrium
Hart, S. and Mas-Colell, A · 2000
Earlier work this paper cites.
Planning in the presence of cost functions controlled by an adversary
McMahan, H. B., Gordon, G. J., and Blum, A · 2003
Earlier work this paper cites.
Dynamic programming for partially observable stochastic games
Hansen, E. A., Bernstein, D. S., and Zilberstein, S · 2004
Earlier work this paper cites.
Regret minimization in games with incomplete information
Zinkevich, M., Johanson, M., Bowling, M., and Piccione, C · 2008
Earlier work this paper cites.
Monte carlo sampling for regret minimization in extensive games
Lanctot, M., Waugh, K., Zinkevich, M., and Bowling, M. H · 2009
Earlier work this paper cites.
An exact double-oracle algorithm for zero-sum extensive-form games with imperfect information
Bosansky, B., Kiekintveld, C., Lisy, V., and Pechoucek, M · 2014
Earlier work this paper cites.
Bounding the support size in extensive form games with imperfect information
Schmid, M., Moravcik, M., and Hladik, M · 2014
Earlier work this paper cites.
Solving large imperfect information games using CFR+
Tammelin, O · 2014
Cited alongside, same era.
Fictitious self-play in extensive-form games
Heinrich, J., Lanctot, M., and Silver, D · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Cited alongside, same era.
Deep reinforcement learning from self-play in imperfect-information games
Heinrich, J. and Silver, D · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D · 2016
Cited alongside, same era.
Double neural counterfactual regret minimization
Li, H., Hu, K., Ge, Z., Jiang, T., Qi, Y., and Song, L · 2018
Later among the works it cites.
RLlib: Abstractions for distributed reinforcement learning
Liang, E., Liaw, R., Nishihara, R., Moritz, P., Fox, R., Goldberg, K., Gonzalez, J., Jordan, M., and Stoica, I · 2018
Later among the works it cites.
Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning
Rashid, T., Samvelyan, M., De Witt, C. S., Farquhar, G., Foerster, J., and Whiteson, S · 2018
Later among the works it cites.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., et al · 2019
Later among the works it cites.
Deep counterfactual regret minimization
Brown, N., Lerer, A., Gross, S., and Sandholm, T · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bansal, T., Pachocki, J., Sidor, S., Sutskever, I., and Mordatch, I · 2017
Cited alongside, same era.
Dynamic thresholding and pruning for regret minimization
Brown, N., Kroer, C., and Sandholm, T · 2017
Cited alongside, same era.
An algorithm for constructing and solving imperfect recall abstractions of large extensive-form games
Čermák, J., Bošansky, B., and Lisy, V · 2017
Cited alongside, same era.
A unified game-theoretic approach to multiagent reinforcement learning
Lanctot, M., Zambaldi, V., Gruslys, A., Lazaridou, A., Tuyls, K., Pérolat, J., Silver, D., and Graepel, T · 2017
Cited alongside, same era.
Multi-agent actor-critic for mixed cooperative-competitive environments
Lowe, R., Wu, Y. I., Tamar, A., Harb, J., Abbeel, O. P., and Mordatch, I · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Counterfactual multi-agent policy gradients
Foerster, J., Farquhar, G., Afouras, T., Nardelli, N., and Whiteson, S · 2018
Cited alongside, same era.
Later among the works it cites.
Human-level performance in 3d multiplayer games with population-based reinforcement learning
Jaderberg, M., Czarnecki, W. M., Dunning, I., Marris, L., Lever, G., Castaneda, A. G., Beattie, C., Rabinowitz, N. C., Morcos, A. S., Ruderman, A., et al · 2019
Later among the works it cites.
Single deep counterfactual regret minimization
Steinberger, E · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Later among the works it cites.
The advantage regret-matching actor-critic
Gruslys, A., Lanctot, M., Munos, R., Timbers, F., Schmid, M., Perolat, J., Morrill, D., Zambaldi, V., Lespiau, J.-B., Schultz, J., et al · 2020
Later among the works it cites.
Evolutionary reinforcement learning for sample-efficient multiagent coordination
Majumdar, S., Khadka, S., Miret, S., Mcaleer, S., and Tumer, K · 2020
Later among the works it cites.
Pipeline psro: A scalable approach for finding approximate nash equilibria in large games
McAleer, S., Lanier, J., Fox, R., and Baldi, P · 2020
Later among the works it cites.
Dream: Deep regret minimization with advantage baselines and model-free learning
Steinberger, E., Lerer, A., and Brown, N · 2020
Later among the works it cites.