Fetching the paper…
Reading the bibliography…
When solving two-player zero-sum games, multi-agent reinforcement learning (MARL) algorithms often create populations of agents where, at each iteration, a new agent is discovered as the best response to a mixture over the opponent population.
Equilibrium points in n-person games
John F Nash et al · 1950
Earlier work this paper cites.
Iterative solution of games by fictitious play
George W Brown · 1951
Earlier work this paper cites.
Theory of games and economic behavior
Oskar Morgenstern and John Von Neumann · 1953
Earlier work this paper cites.
Game theory for applied economists
Robert S Gibbons · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Planning in the presence of cost functions controlled by an adversary
H Brendan McMahan, Geoffrey J Gordon, and Avrim Blum · 2003
Earlier work this paper cites.
Discrete colonel blotto and general lotto games
Sergiu Hart · 2008
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Using response functions to measure strategy strength
Trevor Davis, Neil Burch, and Michael Bowling · 2014
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation, 2015
Jonathan Long, Evan Shelhamer, and Trevor Darrell · 2015
Earlier work this paper cites.
Rl: Fast reinforcement learning via slow reinforcement learning
Yan Duan, John Schulman, Xi Chen, Peter L Bartlett, Ilya Sutskever, and Pieter Abbeel · 2016
Earlier work this paper cites.
9. a simplified two-person poker
Harold W Kuhn · 2016
Earlier work this paper cites.
Learning to reinforcement learn
Jane X Wang, Zeb Kurth-Nelson, Dhruva Tirumala, Hubert Soyer, Joel Z Leibo, Remi Munos, Charles Blundell, Dharshan Kumaran, and Matt Botvinick · 2016
Earlier work this paper cites.
Continuous adaptation via meta-learning in nonstationary and competitive environments
Maruan Al-Shedivat, Trapit Bansal, Yuri Burda, Ilya Sutskever, Igor Mordatch, and Pieter Abbeel · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Learning with opponent-learning awareness
Jakob N Foerster, Richard Y Chen, Maruan Al-Shedivat, Shimon Whiteson, Pieter Abbeel, and Igor Mordatch · 2017
Earlier work this paper cites.
A unified game-theoretic approach to multiagent reinforcement learning, 2017
Marc Lanctot, Vinicius Zambaldi, Audrunas Gruslys, Angeliki Lazaridou, Karl Tuyls, Julien Perolat, David Silver, and Thore Graepel · 2017
Earlier work this paper cites.
Pointnet: Deep learning on point sets for 3d classification and segmentation
Charles R Qi, Hao Su, Kaichun Mo, and Leonidas J Guibas · 2017
Earlier work this paper cites.
Evolution strategies as a scalable alternative to reinforcement learning, 2017
Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Value-decomposition networks for cooperative multi-agent learning
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al · 2017
Earlier work this paper cites.
Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabas Poczos, Ruslan Salakhutdinov, and Alexander Smola · 2017
Cited alongside, same era.
Structured evolution with compact architectures for scalable policy optimization, 2018
Krzysztof Choromanski, Mark Rowland, Vikas Sindhwani, Richard E. Turner, and Adrian Weller · 2018
Cited alongside, same era.
Dice: The infinitely differentiable monte carlo estimator
Jakob Foerster, Gregory Farquhar, Maruan Al-Shedivat, Tim Rocktäschel, Eric Xing, and Shimon Whiteson · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado Van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2018
Cited alongside, same era.
Rein Houthooft, Richard Y Chen, Phillip Isola, Bradly C Stadie, Filip Wolski, Jonathan Ho, and Pieter Abbeel · 2018
Optimizing millions of hyperparameters by implicit differentiation
Jonathan Lorraine, Paul Vicol, and David Duvenaud · 2020
Later among the works it cites.
Pipeline PSRO: A scalable approach for finding approximate nash equilibria in large games
Stephen McAleer, John Lanier, Roy Fox, and Pierre Baldi · 2020
Later among the works it cites.
Discovering reinforcement learning algorithms
Junhyuk Oh, Matteo Hessel, Wojciech M Czarnecki, Zhongwen Xu, Hado van Hasselt, Satinder Singh, and David Silver · 2020
Later among the works it cites.
Parrot: Data-driven behavioral priors for reinforcement learning
Avi Singh, Huihan Liu, Gaoyue Zhou, Albert Yu, Nicholas Rhinehart, and Sergey Levine · 2020
Later among the works it cites.
A deterministic linear program solver in current matrix multiplication time
Jan van den Brand · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Promp: Proximal meta-policy search
Jonas Rothfuss, Dennis Lee, Ignasi Clavera, Tamim Asfour, and Pieter Abbeel · 2018
Cited alongside, same era.
Meta-gradient reinforcement learning
Zhongwen Xu, Hado van Hasselt, and David Silver · 2018
Cited alongside, same era.
On learning intrinsic rewards for policy gradient methods
Zeyu Zheng, Junhyuk Oh, and Satinder Singh · 2018
Cited alongside, same era.
Emergent tool use from multi-agent autocurricula
Bowen Baker, Ingmar Kanitscheider, Todor Markov, Yi Wu, Glenn Powell, Bob McGrew, and Igor Mordatch · 2019
Cited alongside, same era.
Open-ended learning in symmetric zero-sum games
D Balduzzi, M Garnelo, Y Bachrach, W Czarnecki, J Pérolat, M Jaderberg, and T Graepel · 2019
Cited alongside, same era.
Improving generalization in meta reinforcement learning using learned objectives
Louis Kirsch, Sjoerd van Steenkiste, and Juergen Schmidhuber · 2019
Cited alongside, same era.
Openspiel: A framework for reinforcement learning in games
Marc Lanctot, Edward Lockhart, Jean-Baptiste Lespiau, Vinicius Zambaldi, Satyaki Upadhyay, Julien Pérolat, Sriram Srinivasan, Finbarr Timbers, Karl Tuyls, Shayegan Omidshafiei, et al · 2019
Cited alongside, same era.
Meta-gradient reinforcement learning with an objective discovered online
Zhongwen Xu, Hado van Hasselt, Matteo Hessel, Junhyuk Oh, Satinder Singh, and David Silver · 2020
Later among the works it cites.
α \alpha α \alpha -rank: Practically scaling α \alpha -rank through stochastic optimisation
Yaodong Yang, Rasul Tutunov, Phu Sakulwongtana, and Haitham Bou Ammar · 2020
Later among the works it cites.
An overview of multi-agent reinforcement learning from game theoretical perspective
Yaodong Yang and Jun Wang · 2020
Later among the works it cites.
Multi-agent determinantal q-learning
Yaodong Yang, Ying Wen, Jun Wang, Liheng Chen, Kun Shao, David Mguni, and Weinan Zhang · 2020
Later among the works it cites.
A self-tuning actor-critic algorithm
Tom Zahavy, Zhongwen Xu, Vivek Veeriah, Matteo Hessel, Junhyuk Oh, Hado van Hasselt, David Silver, and Satinder Singh · 2020
Later among the works it cites.
What can learned intrinsic rewards capture?
Zeyu Zheng, Junhyuk Oh, Matteo Hessel, Zhongwen Xu, Manuel Kroiss, Hado Van Hasselt, David Silver, and Satinder Singh · 2020
Later among the works it cites.
Online meta-critic learning for off-policy actor-critic methods
Wei Zhou, Yiying Li, Yongxin Yang, Huaimin Wang, and Timothy Hospedales · 2020
Later among the works it cites.
Meta learning via learned loss
Sarah Bechtle, Artem Molchanov, Yevgen Chebotar, Edward Grefenstette, Ludovic Righetti, Gaurav Sukhatme, and Franziska Meier · 2021
Closest in time.
On the complexity of computing markov perfect equilibrium in general-sum stochastic games
Xiaotie Deng, Yuhao Li, David Henry Mguni, Jun Wang, and Yaodong Yang · 2021
Closest in time.
Le Cong Dinh, Yaodong Yang, Zheng Tian, Nicolas Perez Nieves, Oliver Slumbers, David Henry Mguni, and Jun Wang · 2021
Closest in time.
Unifying behavioral and response diversity for open-ended learning in zero-sum games
Xiangyu Liu, Hangtian Jia, Ying Wen, Yaodong Yang, Yujing Hu, Yingfeng Chen, Changjie Fan, and Zhipeng Hu · 2021
Closest in time.
XDO: A double oracle algorithm for extensive-form games
Stephen McAleer, John Lanier, Pierre Baldi, and Roy Fox · 2021
Closest in time.
Modelling behavioural diversity for learning in open-ended games
Nicolas Perez Nieves, Yaodong Yang, Oliver Slumbers, David Henry Mguni, and Jun Wang · 2021
Closest in time.
Measuring the non-transitivity in chess, 2021
Ricky Sanjaya, Jun Wang, and Yaodong Yang · 2021
Closest in time.
Credit assignment with meta-policy gradient for multi-agent reinforcement learning
Jianzhun Shao, Hongchang Zhang, Yuhang Jiang, Shuncheng He, and Xiangyang Ji · 2021
Closest in time.
Open-ended learning leads to generally capable agents
Ended Learning Team, Adam Stooke, Anuj Mahajan, Catarina Barros, Charlie Deck, Jakob Bauer, Jakub Sygnowski, Maja Trebacz, Max Jaderberg, Michael Mathieu, et al · 2021
Closest in time.
Diverse auto-curriculum is critical for successful real-world multiagent learning systems
Yaodong Yang, Jun Luo, Ying Wen, Oliver Slumbers, Daniel Graves, Haitham Bou Ammar, Jun Wang, and Matthew E Taylor · 2021
Closest in time.