Fetching the paper…
Reading the bibliography…
In multi-agent settings with mixed incentives, methods developed for zero-sum games have been shown to lead to detrimental outcomes.
Bayesian action decoder for deep multi-agent reinforcement learning. In International Conference on Machine Learning . PMLR, 1942–1951
Jakob Foerster, Francis Song, Edward Hughes, Neil Burch, Iain Dunning, Shimon Whiteson, Matthew Botvinick, and Michael Bowling. 2019 · 1951
Earlier work this paper cites.
Stochastic Games
L. S. Shapley. 1953 · 1953
Earlier work this paper cites.
" Prisoner’s Dilemma" and" Chicken" Models in International Politics
Glenn H Snyder. 1971 · 1971
Earlier work this paper cites.
Social dilemmas
Robyn M Dawes. 1980 · 1980
Earlier work this paper cites.
The evolution of cooperation
Robert Axelrod and William D Hamilton. 1981 · 1981
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Iterated Prisoner’s Dilemma contains strategies that dominate any evolutionary opponent
William H. Press and Freeman J. Dyson. 2012 · 2012
Earlier work this paper cites.
RL$ˆ2$: Fast Reinforcement Learning via Slow Reinforcement Learning
Yan Duan, John Schulman, Xi Chen, Peter L. Bartlett, Ilya Sutskever, and Pieter Abbeel. 2016 · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Vedavyas Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy P. Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis. 2016 · 2016
Earlier work this paper cites.
Libratus: the superhuman AI for no-limit poker. In Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence
Noam Brown and Tuomas Sandholm. 2017 · 2017
Earlier work this paper cites.
Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks. In International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 70) . 1126–1135
Chelsea Finn, Pieter Abbeel, and Sergey Levine. 2017 · 2017
Earlier work this paper cites.
Reinforcement learning produces dominant strategies for the Iterated Prisoner’s Dilemma
Marc Harper, Vincent Knight, Martin Jones, Georgios Koutsovoulos, Nikoleta E. Glynatsi, and Owen Campbell. 2017 · 2017
Earlier work this paper cites.
Maintaining cooperation in complex social dilemmas using deep reinforcement learning
Adam Lerer and Alexander Peysakhovich. 2017 · 2017
Earlier work this paper cites.
Evolution Strategies as a Scalable Alternative to Reinforcement Learning
Tim Salimans, Jonathan Ho, Xi Chen, and Ilya Sutskever. 2017 · 2017
Earlier work this paper cites.
Proximal Policy Optimization Algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Cited alongside, same era.
Continuous Adaptation via Meta-Learning in Nonstationary and Competitive Environments. In International Conference on Learning Representations
Maruan Al-Shedivat, Trapit Bansal, Yura Burda, Ilya Sutskever, Igor Mordatch, and Pieter Abbeel. 2018 · 2018
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang. 2018 · 2018
Cited alongside, same era.
Learning with Opponent-Learning Awareness. In Proceedings of the 17th International Conference on Autonomous Agents and MultiAgent Systems . 122–130
Jakob Foerster, Richard Y Chen, Maruan Al-Shedivat, Shimon Whiteson, Pieter Abbeel, and Igor Mordatch. 2018 · 2018
Cited alongside, same era.
Open Problems in Cooperative AI. In Cooperative AI workshop
Allan Dafoe, Edward Hughes, Yoram Bachrach, Tantum Collins, Kevin R McKee, Joel Z Leibo, Kate Larson, and Thore Graepel. 2021 · 2021
Later among the works it cites.
Scalable Evaluation of Multi-Agent Reinforcement Learning with Melting Pot. In Proceedings of the 38th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 139) . PMLR, 6187–6199
Joel Z. Leibo, Edgar A. Duéñez-Guzmán, Alexander Vezhnevets, John P. Agapiou, Peter Sunehag, Raphael Koster, Jayd Matyas, Charlie Beattie, Igor Mordatch, and Thore Graepel. 2021 · 2021
Later among the works it cites.
L2E: Learning to Exploit Your Opponent
Zhe Wu, Kai Li, Enmin Zhao, Hang Xu, Meng Zhang, Haobo Fu, Bo An, and Junliang Xing. 2021 · 2021
Later among the works it cites.
Human-level play in the game of Diplomacy by combining language models with strategic reasoning
Anton Bakhtin, Noam Brown, Emily Dinan, Gabriele Farina, Colin Flaherty, Daniel Fried, Andrew Goff, Jonathan Gray, Hengyuan Hu, Athul Paul Jacob, Mojtaba Komeili, Karthik Konath, Minae Kwon, Adam Lerer, Mike Lewis, Alexander H. Miller, Sasha Mitts, Adithya Renduchintala, Stephen Roller, Dirk Rowe, Weiyan Shi, Joe Spisak, Alexander Wei, David Wu, Hugh Zhang, and Markus Zijlstra. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning. In International conference on machine learning . PMLR, 4295–4304
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson. 2018 · 2018
Cited alongside, same era.
Human-level performance in 3D multiplayer games with population-based reinforcement learning
Max Jaderberg, Wojciech M. Czarnecki, Iain Dunning, Luke Marris, Guy Lever, Antonio Garcia Castañeda, Charles Beattie, Neil C. Rabinowitz, Ari S. Morcos, Avraham Ruderman, Nicolas Sonnerat, Tim Green, Louise Deason, Joel Z. Leibo, David Silver, Demis Hassabis, Koray Kavukcuoglu, and Thore Graepel. 2019 · 2019
Cited alongside, same era.
Social influence as intrinsic motivation for multi-agent deep reinforcement learning. In International conference on machine learning . PMLR, 3040–3049
Natasha Jaques, Angeliki Lazaridou, Edward Hughes, Caglar Gulcehre, Pedro Ortega, DJ Strouse, Joel Z Leibo, and Nando De Freitas. 2019 · 2019
Cited alongside, same era.
Differentiable Game Mechanics
Alistair Letcher, David Balduzzi, Sébastien Racanière, James Martens, Jakob N. Foerster, Karl Tuyls, and Thore Graepel. 2019a · 2019
Cited alongside, same era.
The StarCraft Multi-Agent Challenge
Mikayel Samvelyan, Tabish Rashid, Christian Schroeder de Witt, Gregory Farquhar, Nantas Nardelli, Tim G. J. Rudner, Chia-Man Hung, Philiph H. S. Torr, Jakob Foerster, and Shimon Whiteson. 2019 · 2019
Cited alongside, same era.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M. Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H. Choi, Richard Powell, Timo Ewalds, Petko Georgiev, Junhyuk Oh, Dan Horgan, Manuel Kroiss, Ivo Danihelka, Aja Huang, Laurent Sifre, Trevor Cai, John P. Agapiou, Max Jaderberg, Alexander Sasha Vezhnevets, Rémi Leblond, Tobias Pohlen, Valentin Dalibard, David Budden, Yury Sulsky, James Molloy, Tom Le Paine, Çaglar Gülçehre, Ziyu Wang, Tobias Pfaff, Yuhuai Wu, Roman Ring, Dani Yogatama, Dario Wünsch, Katrina McKinney, Oliver Smith, Tom Schaul, Timothy P. Lillicrap, Koray Kavukcuoglu, Demis Hassabis, Chris Apps, and David Silver. 2019 · 2019
Cited alongside, same era.
Haiku: Sonnet for JAX
Tom Hennigan, Trevor Cai, Tamara Norman, and Igor Babuschkin. 2020 · 2020
Cited alongside, same era.
Model-free conventions in multi-agent reinforcement learning with heterogeneous preferences
Raphael Köster, Kevin R McKee, Richard Everett, Laura Weidinger, William S Isaac, Edward Hughes, Edgar A Duéñez-Guzmán, Thore Graepel, Matthew Botvinick, and Joel Z Leibo. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
The Good Shepherd: An Oracle Agent for Mechanism Design
Jan Balaguer, Raphael Koster, Christopher Summerfield, and Andrea Tacchetti. 2022 · 2022
Later among the works it cites.
SMACv2: An Improved Benchmark for Cooperative Multi-Agent Reinforcement Learning
Benjamin Ellis, Skander Moalla, Mikayel Samvelyan, Mingfei Sun, Anuj Mahajan, Jakob N. Foerster, and Shimon Whiteson. 2022 · 2022
Later among the works it cites.
Discovered Policy Optimisation
Chris Lu, Jakub Grudzien Kuba, Alistair Letcher, Luke Metz, Christian Schröder de Witt, and Jakob N. Foerster. 2022a · 2022
Later among the works it cites.
Adversarial Cheap Talk. In Decision Awareness in Reinforcement Learning Workshop at ICML 2022
Chris Lu, Timon Willi, Alistair Letcher, and Jakob Nicolaus Foerster. 2022c · 2022
Later among the works it cites.
COLA: Consistent Learning with Opponent-Learning Awareness. In International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 162) . 23804–23831
Timon Willi, Alistair Letcher, Johannes Treutlein, and Jakob N. Foerster. 2022 · 2022
Later among the works it cites.
Adherence Improves Cooperation in Sequential Social Dilemmas
Yuyu Yuan, Ting Guo, Pengqian Zhao, and Hongpu Jiang. 2022 · 2022
Later among the works it cites.
Proximal Learning With Opponent-Learning Awareness
Stephen Zhao, Chris Lu, Roger Baker Grosse, and Jakob Nicolaus Foerster. 2022 · 2022
Later among the works it cites.
Analyzing the Sample Complexity of Model-Free Opponent Shaping. In ICML Workshop on New Frontiers in Learning, Control, and Dynamical Systems
Kitty Fung, Qizhen Zhang, Chris Lu, Timon Willi, and Jakob Nicolaus Foerster. 2023 · 2023
Closest in time.
Pax: Multi-Agent Learning in JAX
Timon Willi, Akbir Khan, Newton Kwan, Mikayel Samvelyan, Chris Lu, and Jakob Foerster. 2023 · 2023
Closest in time.