Fetching the paper…
Reading the bibliography…
In this report, we present results reproductions for several core algorithms implemented in the OpenSpiel framework for learning in games.
Policy gradient methods for reinforcement learning with function approximation
Richard Sutton, David Mcallester, Satinder Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Planning in the presence of cost functions controlled by an adversary
H. Brendan McMahan, Geoffrey J. Gordon, and Avrim Blum · 2003
Earlier work this paper cites.
Generalised weakened fictitious play
David S. Leslie and E.J. Collins · 2006
Earlier work this paper cites.
Solving games with functional regret estimation, 2014
Kevin Waugh, Dustin Morrill, J. Andrew Bagnell, and Michael Bowling · 2014
Earlier work this paper cites.
Fictitious self-play in extensive-form games
Johannes Heinrich, Marc Lanctot, and David Silver · 2015
Earlier work this paper cites.
Contributions to the Theory of Games (AM-24), Volume I
H.W. Kuhn and A.W. Tucker · 2016
Earlier work this paper cites.
Deep reinforcement learning from self-play in imperfect-information games
Johannes Heinrich and David Silver · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Timothy P. Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Efficient hyperparameter optimization and infinitely many armed bandits
Lisha Li, Kevin G. Jamieson, Giulia DeSalvo, Afshin Rostamizadeh, and Ameet Talwalkar · 2016
Cited alongside, same era.
A unified game-theoretic approach to multiagent reinforcement learning, 2017
Marc Lanctot, Vinicius Zambaldi, Audrunas Gruslys, Angeliki Lazaridou, Karl Tuyls, Julien Perolat, David Silver, and Thore Graepel · 2017
Cited alongside, same era.
Actor-critic policy optimization in partially observable multiagent environments
Sriram Srinivasan, Marc Lanctot, Vinícius Flores Zambaldi, Julien Pérolat, Karl Tuyls, Rémi Munos, and Michael Bowling · 2018
Cited alongside, same era.
Mean actor critic, 2018
Cameron Allen, Kavosh Asadi, Melrose Roderick, Abdel rahman Mohamed, George Konidaris, and Michael Littman · 2018
Cited alongside, same era.
Computing approximate equilibria in sequential adversarial games by exploitability descent
Edward Lockhart, Marc Lanctot, Julien Pérolat, Jean-Baptiste Lespiau, Dustin Morrill, Finbarr Timbers, and Karl Tuyls · 2019
Later among the works it cites.
Single deep counterfactual regret minimization
Eric Steinberger · 2019
Later among the works it cites.
A generalized training approach for multiagent learning
Paul Muller, Shayegan Omidshafiei, Mark Rowland, Karl Tuyls, Julien Pérolat, Siqi Liu, Daniel Hennes, Luke Marris, Marc Lanctot, Edward Hughes, Zhe Wang, Guy Lever, Nicolas Heess, Thore Graepel, and Rémi Munos · 2019
Later among the works it cites.
Experiment tracking with weights and biases, 2020
Lukas Biewald · 2020
Later among the works it cites.
An overview of multi-agent reinforcement learning from game theoretical perspective, 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Marc Lanctot, Edward Lockhart, Jean-Baptiste Lespiau, Vinicius Zambaldi, Satyaki Upadhyay, Julien Pérolat, Sriram Srinivasan, Finbarr Timbers, Karl Tuyls, Shayegan Omidshafiei, Daniel Hennes, Dustin Morrill, Paul Muller, Timo Ewalds, Ryan Faulkner, János Kramár, Bart De Vylder, Brennan Saeta, James Bradbury, David Ding, Sebastian Borgeaud, Matthew Lai, Julian Schrittwieser, Thomas Anthony, Edward Hughes, Ivo Danihelka, and Jonah Ryan-Davis · 2019
Cited alongside, same era.
A survey and critique of multiagent deep reinforcement learning
Pablo Hernandez-Leal, Bilal Kartal, and Matthew E. Taylor · 2019
Cited alongside, same era.
Deep counterfactual regret minimization, 2019
Noam Brown, Adam Lerer, Sam Gross, and Tuomas Sandholm · 2019
Cited alongside, same era.
Yaodong Yang and Jun Wang · 2020
Later among the works it cites.
Neural replicator dynamics, 2020
Daniel Hennes, Dustin Morrill, Shayegan Omidshafiei, Remi Munos, Julien Perolat, Marc Lanctot, Audrunas Gruslys, Jean-Baptiste Lespiau, Paavo Parmas, Edgar Duenez-Guzman, and Karl Tuyls · 2020
Later among the works it cites.
The advantage regret-matching actor-critic
Audrūnas Gruslys, Marc Lanctot, Remi Munos, Finbarr Timbers, Martin Schmid, Pérolat Julien, Dustin Morrill, Vinicius Zambaldi, Jean-Baptiste Lespiau, John Schultz, Mohammad Azar, Michael Bowling, and Karl Tuyls · 2020
Later among the works it cites.