Fetching the paper…
Reading the bibliography…
In various real-world scenarios, interactions among agents often resemble the dynamics of general-sum games, where each agent strives to optimize its own utility.
On the theory of games of strategy
John von Neumann · 1928
Earlier work this paper cites.
Stochastic games
Lloyd Shapley · 1953
Earlier work this paper cites.
Effective choice in the prisoner’s dilemma
Robert Axelrod · 1980
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald Williams · 1992
Earlier work this paper cites.
Actor-critic algorithms, 2000
Vijay R. Konda and John N. Tsitsiklis · 2000
Earlier work this paper cites.
Playing atari with deep reinforcement learning, 2013
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Continuous adaptation via meta-learning in nonstationary and competitive environments, 2018
Maruan Al-Shedivat, Trapit Bansal, Yuri Burda, Ilya Sutskever, Igor Mordatch, and Pieter Abbeel · 2018
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor, 2018
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Maintaining cooperation in complex social dilemmas using deep reinforcement learning, 2018
Adam Lerer and Alexander Peysakhovich · 2018
Cited alongside, same era.
Loaded dice: Trading off bias and variance in any-order score function estimators for reinforcement learning, 2019
Gregory Farquhar, Shimon Whiteson, and Jakob Foerster · 2019
Cited alongside, same era.
Stable opponent shaping in differentiable games, 2021
Alistair Letcher, Jakob Foerster, David Balduzzi, Tim Rocktäschel, and Shimon Whiteson · 2021
Later among the works it cites.
Model-free opponent shaping, 2022
Chris Lu, Timon Willi, Christian Schroeder de Witt, and Jakob Foerster · 2022
Later among the works it cites.
Cola: Consistent learning with opponent-learning awareness, 2022
Timon Willi, Alistair Letcher, Johannes Treutlein, and Jakob Foerster · 2022
Later among the works it cites.
Proximal learning with opponent-learning awareness, 2022
Stephen Zhao, Chris Lu, Roger Baker Grosse, and Jakob Nicolaus Foerster · 2022
Later among the works it cites.
Meta-value learning: a general framework for learning with learning awareness, 2023
Tim Cooijmans, Milad Aghajohari, and Aaron Courville · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforcement Learning: Theory and Algorithms
Alekh Agarwal, Nan Jiang, Sham Kakade, and Wen Sun · 2021
Cited alongside, same era.
A policy gradient algorithm for learning to learn in multiagent reinforcement learning, 2021
Dong-Ki Kim, Miao Liu, Matthew Riemer, Chuangchuang Sun, Marwa Abdulhai, Golnaz Habibi, Sebastian Lopez-Cot, Gerald Tesauro, and Jonathan P. How · 2021
Cited alongside, same era.
Dice: The infinitely differentiable monte-carlo estimator, 2018a
Jakob Foerster, Gregory Farquhar, Maruan Al-Shedivat, Tim Rocktäschel, Eric P. Xing, and Shimon Whiteson
Cited in the paper.
Learning with opponent-learning awareness, 2018b
Jakob N. Foerster, Richard Y. Chen, Maruan Al-Shedivat, Shimon Whiteson, Pieter Abbeel, and Igor Mordatch
Cited in the paper.
Milad Aghajohari, Tim Cooijmans, Juan Agustin Duque, Shunichi Akatsuka, and Aaron Courville · 2024
Closest in time.