Fetching the paper…
Reading the bibliography…
This work studies an algorithm, which we call magnetic mirror descent, that is inspired by mirror descent and the non-Euclidean proximal gradient algorithm.
OpenSpiel: A framework for reinforcement learning in games
Marc Lanctot, Edward Lockhart, Jean-Baptiste Lespiau, Vinicius Zambaldi, Satyaki Upadhyay, Julien Pérolat, Sriram Srinivasan, Finbarr Timbers, Karl Tuyls, Shayegan Omidshafiei, Daniel Hennes, Dustin Morrill, Paul Muller, Timo Ewalds, Ryan Faulkner, János Kramár, Bart De Vylder, Brennan Saeta, James Bradbury, David Ding, Sebastian Borgeaud, Matthew Lai, Julian Schrittwieser, Thomas Anthony, Edward Hughes, Ivo Danihelka, and Jonah Ryan-Davis · 1908
Earlier work this paper cites.
Theory of games and economic behavior
J. von Neumann and O. Morgenstern · 1947
Earlier work this paper cites.
Iterative solutions of games by fictitious play
G.W. Brown · 1951
Earlier work this paper cites.
9. a simplified two-person poker
Helmut Kuhn · 1951
Earlier work this paper cites.
The second scientific american book of mathematical puzzles and diversions
Aaron Bakst and Martin Gardner · 1962
Earlier work this paper cites.
Reduction of a game with full memory to a matrix game
IV Romanovskii · 1962
Earlier work this paper cites.
Monotone operators associated with saddle.functions and minimax problems, 1970
R. Tyrrell Rockafellar · 1970
Earlier work this paper cites.
The extragradient method for finding saddle points and other problems
G. M. Korpelevich · 1976
Earlier work this paper cites.
A modification of the Arrow-Hurwicz method for search of saddle points
Leonid Denisovich Popov · 1980
Earlier work this paper cites.
Problem Complexity and Method Efficiency in Optimization
A.S. Nemirovsky and D.B. Yudin · 1983
Earlier work this paper cites.
Models for the Game of Liar’s Dice , pp. 15–28
Christopher P. Ferguson and Thomas S. Ferguson · 1991
Earlier work this paper cites.
Quantal response equilibria for normal form games
Richard McKelvey and Thomas Palfrey · 1995
Earlier work this paper cites.
Efficient computation of equilibria for extensive two-person games
Daphne Koller, Nimrod Megiddo, and Bernhard Von Stengel · 1996
Earlier work this paper cites.
Efficient computation of behavior strategies
Bernhard Von Stengel · 1996
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Leslie Pack Kaelbling, Michael L. Littman, and Anthony R. Cassandra · 1998
Earlier work this paper cites.
Quantal response equilibria for extensive form games
Richard D. McKelvey and Thomas R. Palfrey · 1998
Earlier work this paper cites.
Bregman monotone optimization algorithms
Heinz H Bauschke, Jonathan M Borwein, and Patrick L Combettes · 2003
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Amir Beck and Marc Teboulle · 2003
Earlier work this paper cites.
Finite-dimensional variational inequalities and complementarity problems
Francisco Facchinei and Jong-Shi Pang · 2003
Earlier work this paper cites.
Planning in the presence of cost functions controlled by an adversary
H. Brendan McMahan, Geoffrey J. Gordon, and Avrim Blum · 2003
Earlier work this paper cites.
Dynamic programming for partially observable stochastic games
Eric A. Hansen, Daniel S. Bernstein, and Shlomo Zilberstein · 2004
Earlier work this paper cites.
Prox-method with rate of convergence o (1/t) for variational inequalities with lipschitz continuous monotone operators and smooth convex-concave saddle point problems
Arkadi Nemirovski · 2004
Earlier work this paper cites.
A dynamic homotopy interpretation of the logistic quantal response equilibrium correspondence
Theodore L. Turocy · 2004
Earlier work this paper cites.
Bayes’ bluff: Opponent modelling in poker
Finnegan Southey, Michael Bowling, Bryce Larson, Carmelo Piccione, Neil Burch, Darse Billings, and Chris Rayner · 2005
Earlier work this paper cites.
Mirror descent policy optimization, 2020
Manan Tomar, Lior Shani, Yonathan Efroni, and Mohammad Ghavamzadeh · 2005
Earlier work this paper cites.
DREAM: deep regret minimization with advantage baselines and model-free learning
Eric Steinberger, Adam Lerer, and Noam Brown · 2006
Earlier work this paper cites.
Algorithmic Game Theory
Noam Nisan, Tim Roughgarden, Eva Tardos, and Vijay V Vazirani (eds.) · 2007
Earlier work this paper cites.
Regret minimization in games with incomplete information
Martin Zinkevich, Michael Johanson, Michael Bowling, and Carmelo Piccione · 2007
Cited alongside, same era.
Multiagent Systems: Algorithmic, Game-Theoretic, and Logical Foundations
Yoav Shoham and Kevin Leyton-Brown · 2008
Cited alongside, same era.
On accelerated proximal gradient methods for convex-concave optimization
Paul Tseng · 2008
Cited alongside, same era.
Maximum entropy inverse reinforcement learning
Brian D. Ziebart, Andrew Maas, J. Andrew Bagnell, and Anind K. Dey · 2008
Cited alongside, same era.
Revisiting design choices in proximal policy optimization
Chloe Ching-Yun Hsu, Celestine Mendler-Dünner, and Moritz Hardt · 2009
Cited alongside, same era.
RLlib: Abstractions for distributed reinforcement learning
Eric Liang, Richard Liaw, Robert Nishihara, Philipp Moritz, Roy Fox, Ken Goldberg, Joseph Gonzalez, Michael Jordan, and Ion Stoica · 2018
Later among the works it cites.
What game are we playing? end-to-end learning in normal and extensive form games
Chun Kai Ling, Fei Fang, and J Zico Kolter · 2018
Later among the works it cites.
Relatively smooth convex optimization by first-order methods, and applications
Haihao Lu, Robert M Freund, and Yurii Nesterov · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Deep counterfactual regret minimization
Noam Brown, Adam Lerer, Sam Gross, and Tuomas Sandholm · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Monte carlo sampling for regret minimization in extensive games
Marc Lanctot, Kevin Waugh, Martin Zinkevich, and Michael Bowling · 2009
Cited alongside, same era.
Composite objective mirror descent
John C Duchi, Shai Shalev-Shwartz, Yoram Singer, and Ambuj Tewari · 2010
Cited alongside, same era.
Smoothing techniques for computing nash equilibria of sequential games
Samid Hoda, Andrew Gilpin, Javier Pena, and Tuomas Sandholm · 2010
Cited alongside, same era.
Approximation accuracy, gradient methods, and error bound for structured convex optimization
Paul Tseng · 2010
Cited alongside, same era.
Computing sequential equilibria using agent quantal response equilibria
Theodore L. Turocy · 2010
Cited alongside, same era.
Convex analysis and monotone operator theory in Hilbert spaces , volume 408
Heinz H Bauschke, Patrick L Combettes, et al · 2011
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
Gabriele Farina, Christian Kroer, and Tuomas Sandholm · 2019
Later among the works it cites.
Differentiable game mechanics
Alistair Letcher, David Balduzzi, Sébastien Racanière, James Martens, Jakob Foerster, Karl Tuyls, and Thore Graepel · 2019
Later among the works it cites.
Large scale learning of agent rationality in two-player zero-sum games
Chun Kai Ling, Fei Fang, and J Zico Kolter · 2019
Later among the works it cites.
No-press diplomacy: Modeling multi-agent gameplay
Philip Paquette, Yuchen Lu, Seton Steven Bocco, Max Smith, Satya O-G, Jonathan K Kummerfeld, Joelle Pineau, Satinder Singh, and Aaron C Courville · 2019
Later among the works it cites.
Variance reduction in monte carlo counterfactual regret minimization (vr-mccfr) for extensive form games using baselines
Martin Schmid, Neil Burch, Marc Lanctot, Matej Moravcik, Rudolf Kadlec, and Michael Bowling · 2019
Later among the works it cites.
Agent57: Outperforming the atari human benchmark
Adrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann, Alex Vitvitskyi, Daniel Guo, and Charles Blundell · 2020
Later among the works it cites.
Low-variance and zero-variance baselines for extensive-form games
Trevor Davis, Martin Schmid, and Michael Bowling · 2020
Later among the works it cites.
Faster algorithms for extensive-form game solving via improved smoothing functions
Christian Kroer, Kevin Waugh, Fatma Kılınç-Karzan, and Tuomas Sandholm · 2020
Later among the works it cites.
Double neural counterfactual regret minimization
Hui Li, Kailiang Hu, Shaohua Zhang, Yuan Qi, and Le Song · 2020
Later among the works it cites.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps
Lior Shani, Yonathan Efroni, and Shie Mannor · 2020
Later among the works it cites.
Leverage the average: an analysis of kl regularization in reinforcement learning
Nino Vieillard, Tadashi Kozuno, Bruno Scherrer, Olivier Pietquin, Remi Munos, and Matthieu Geist · 2020
Later among the works it cites.
No-press diplomacy from scratch
Anton Bakhtin, David Wu, Adam Lerer, and Noam Brown · 2021
Later among the works it cites.
Fast policy extragradient methods for competitive games with entropy regularization
Shicong Cen, Yuting Wei, and Yuejie Chi · 2021
Later among the works it cites.
The advantage regret-matching actor-critic, 2021
Audrunas Gruslys, Marc Lanctot, Remi Munos, Finbarr Timbers, Martin Schmid, Julien Perolet, Dustin Morrill, Vinicius Zambaldi, Jean-Baptiste Lespiau, John Schultz, Mohammad Gheshlaghi Azar, Michael Bowling, and Karl Tuyls · 2021
Later among the works it cites.
Accelerated bregman proximal gradient methods for relatively smooth convex optimization
Filip Hanzely, Peter Richtarik, and Lin Xiao · 2021
Later among the works it cites.
XDO: A double oracle algorithm for extensive-form games
Stephen Marcus McAleer, John Banister Lanier, Kevin Wang, Pierre Baldi, and Roy Fox · 2021
Later among the works it cites.
From poincaré recurrence to convergence in imperfect information games: Finding equilibrium via regularization
Julien Pérolat, Rémi Munos, Jean-Baptiste Lespiau, Shayegan Omidshafiei, Mark Rowland, Pedro A. Ortega, Neil Burch, Thomas W. Anthony, David Balduzzi, Bart De Vylder, Georgios Piliouras, Marc Lanctot, and Karl Tuyls · 2021
Later among the works it cites.
Generalized data distribution iteration
Jiajun Fan and Changnan Xiao · 2022
Closest in time.
Extragradient method: O(1/k) last-iterate convergence for monotone variational inequalities and connections with cocoercivity
Eduard Gorbunov, Nicolas Loizou, and Gauthier Gidel · 2022
Closest in time.
The 37 implementation details of proximal policy optimization
Shengyi Huang, Rousslan Fernand Julien Dossa, Antonin Raffin, Anssi Kanervisto, and Weixun Wang · 2022
Closest in time.
Modeling strong and human-like gameplay with KL-regularized search
Athul Paul Jacob, David J Wu, Gabriele Farina, Adam Lerer, Hengyuan Hu, Anton Bakhtin, Jacob Andreas, and Noam Brown · 2022
Closest in time.
Equilibrium finding in normal-form games via greedy regret minimization, 2022
Hugh Zhang, Adam Lerer, and Noam Brown · 2022
Closest in time.