Fetching the paper…
Reading the bibliography…
Learning With Opponent-Learning Awareness (LOLA) (Foerster et al.
Social dilemmas
Robyn M Dawes · 1980
Earlier work this paper cites.
The evolution of cooperation
Robert Axelrod and William Donald Hamilton · 1981
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
Multi-agent learning with policy prediction
Chongjie Zhang and Victor Lesser · 2010
Earlier work this paper cites.
No-regret reductions for imitation learning and structured prediction
Stéphane Ross, Geoffrey J Gordon, and J Andrew Bagnell · 2011
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
Kyunghyun Cho, Bart Van Merriënboer, Dzmitry Bahdanau, and Yoshua Bengio · 2014
Earlier work this paper cites.
Proximal algorithms
Neal Parikh and Stephen Boyd · 2014
Earlier work this paper cites.
Maintaining cooperation in complex social dilemmas using deep reinforcement learning
Adam Lerer and Alexander Peysakhovich · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Continuous adaptation via meta-learning in nonstationary and competitive environments
Maruan Al-Shedivat, Trapit Bansal, Yura Burda, Ilya Sutskever, Igor Mordatch, and Pieter Abbeel · 2018
Cited alongside, same era.
Inequity aversion improves cooperation in intertemporal social dilemmas
Edward Hughes, Joel Z Leibo, Matthew Phillips, Karl Tuyls, Edgar Dueñez-Guzman, Antonio García Castañeda, Iain Dunning, Tina Zhu, Kevin McKee, Raphael Koster, et al · 2018
Cited alongside, same era.
Stable opponent shaping in differentiable games
Alistair Letcher, Jakob Foerster, David Balduzzi, Tim Rocktäschel, and Shimon Whiteson · 2018
Cited alongside, same era.
Evolving intrinsic motivations for altruistic behavior
Jane X Wang, Edward Hughes, Chrisantha Fernando, Wojciech M Czarnecki, Edgar A Duéñez-Guzmán, and Joel Z Leibo · 2018
Cited alongside, same era.
Loaded dice: trading off bias and variance in any-order score function estimators for reinforcement learning
A unified analysis of extra-gradient and optimistic gradient methods for saddle point problems: Proximal point approach
Aryan Mokhtari, Asuman Ozdaglar, and Sarath Pattathil · 2020
Later among the works it cites.
Status-quo policy gradient in multi-agent reinforcement learning
Pinkesh Badjatiya, Mausoom Sarkar, Nikaash Puri, Jayakumar Subramanian, Abhishek Sinha, Siddharth Singh, and Balaji Krishnamurthy · 2021
Later among the works it cites.
Minimax optimization with smooth algorithmic adversaries
Tanner Fiez, Chi Jin, Praneeth Netrapalli, and Lillian J Ratliff · 2021
Later among the works it cites.
A game-theoretic approach to multi-agent trust region optimization
Ying Wen, Hui Chen, Yaodong Yang, Zheng Tian, Minne Li, Xu Chen, and Jun Wang · 2021
Later among the works it cites.
Provably efficient policy gradient methods for two-player zero-sum markov games
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gregory Farquhar, Shimon Whiteson, and Jakob Foerster · 2019
Cited alongside, same era.
Social influence as intrinsic motivation for multi-agent deep reinforcement learning
Natasha Jaques, Angeliki Lazaridou, Edward Hughes, Caglar Gulcehre, Pedro Ortega, DJ Strouse, Joel Z Leibo, and Nando De Freitas · 2019
Cited alongside, same era.
Emergence of cooperation in n-person dilemmas through actor-critic reinforcement learning
João Vitor Barbosa, Anna H Reali Costa, Francisco S Melo, Jaime S Sichman, and Francisco C Santos · 2020
Cited alongside, same era.
Social diversity and social preferences in mixed-motive reinforcement learning
Kevin R McKee, Ian Gemp, Brian McWilliams, Edgar A Duéñez-Guzmán, Edward Hughes, and Joel Z Leibo · 2020
Cited alongside, same era.
Learning with opponent-learning awareness
Jakob Foerster, Richard Y Chen, Maruan Al-Shedivat, Shimon Whiteson, Pieter Abbeel, and Igor Mordatch
Cited in the paper.
Dice: The infinitely differentiable monte carlo estimator
Jakob Foerster, Gregory Farquhar, Maruan Al-Shedivat, Tim Rocktäschel, Eric Xing, and Shimon Whiteson
Cited in the paper.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz
Cited in the paper.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel
Cited in the paper.
Yulai Zhao, Yuandong Tian, Jason D Lee, and Simon S Du · 2021
Later among the works it cites.
Mirror learning: A unifying framework of policy optimisation
Jakub Grudzien Kuba, Christian Schroeder de Witt, and Jakob Foerster · 2022
Closest in time.
Chris Lu, Timon Willi, Christian Schroeder de Witt, and Jakob Foerster · 2022
Closest in time.
Cola: Consistent learning with opponent-learning awareness
Timon Willi, Johannes Treutlein, Alistair Letcher, and Jakob Foerster · 2022
Closest in time.