Fetching the paper…
Reading the bibliography…
Regret minimization has played a key role in online learning, equilibrium computation in games, and reinforcement learning (RL).
OpenSpiel: A framework for reinforcement learning in games
Marc Lanctot, Edward Lockhart, Jean-Baptiste Lespiau, Vinicius Zambaldi, Satyaki Upadhyay, Julien Pérolat, Sriram Srinivasan, and Finbarr Timbers et al · 1908
Earlier work this paper cites.
A simple adaptive procedure leading to correlated equilibrium
S. Hart and A. Mas-Colell · 2000
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
D. Precup, R.S. Sutton, and S. Singh · 2000
Earlier work this paper cites.
Experts in a markov decision process
Eyal Even-Dar, Sham M Kakade, and Yishay Mansour · 2005
Earlier work this paper cites.
Learning, regret minimization, and equilibria
A. Blum and Y. Mansour · 2007
Earlier work this paper cites.
Regret minimization in games with incomplete information
M. Zinkevich, M. Johanson, M. Bowling, and C. Piccione · 2008
Earlier work this paper cites.
Monte Carlo sampling for regret minimization in extensive games
Marc Lanctot, Kevin Waugh, Martin Zinkevich, and Michael Bowling · 2009
Earlier work this paper cites.
Multiagent Systems: Algorithmic, Game-Theoretic, and Logical Foundations
Y. Shoham and K. Leyton-Brown · 2009
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
S. Bubeck and N. Cesa-Bianchi · 2012
Earlier work this paper cites.
Tractable objectives for robust policy optimization
Katherine Chen and Michael Bowling · 2012
Earlier work this paper cites.
Efficient monte carlo counterfactual regret minimization in games with many player actions
Richard Gibson, Neil Burch, Marc Lanctot, and Duane Szafron · 2012
Earlier work this paper cites.
Solving imperfect information games using decomposition
Neil Burch, Michael Johanson, and Michael Bowling · 2014
Earlier work this paper cites.
Heads-up Limit Hold’em Poker is solved
Michael Bowling, Neil Burch, Michael Johanson, and Oskari Tammelin · 2015
Earlier work this paper cites.
Fictitious self-play in extensive-form games
Johannes Heinrich, Marc Lanctot, and David Silver · 2015
Cited alongside, same era.
Solving games with functional regret estimation
Kevin Waugh, Dustin Morrill, J. Andrew Bagnell, and Michael Bowling · 2015
Cited alongside, same era.
Deep reinforcement learning from self-play in imperfect-information games
Johannes Heinrich and David Silver · 2016
Cited alongside, same era.
Eqilibrium approximation quality of current no-limit poker bots
Viliam Lisý and Michael H. Bowling · 2016
Cited alongside, same era.
Understanding and improving convolutional neural networks via concatenated rectified linear units
Wenling Shang, Kihyuk Sohn, Diogo Almeida, and Honglak Lee · 2016
Cited alongside, same era.
Reinforcement Learning: An Introduction
R. Sutton and A. Barto · 2018
Later among the works it cites.
Politex: Regret bounds for policy iteration using expert prediction
Yasin Abbasi-Yadkori, Peter Bartlett, Kush Bhatia, Nevena Lazic, Csaba Szepesvari, and Gellert Weisz · 2019
Later among the works it cites.
Superhuman AI for multiplayer poker
Noam Brown and Tuomas Sandholm · 2019
Later among the works it cites.
Learning to correlate in multi-player general-sum sequential games, 2019
Andrea Celli, Alberto Marchesi, Tommaso Bianchi, and Nicola Gatti · 2019
Later among the works it cites.
Coarse correlation in extensive-form games, 2019
Gabriele Farina, Tommaso Bianchi, and Tuomas Sandholm · 2019
Later among the works it cites.
Rethinking formal models of partially observable multiagent decision making
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Superhuman AI for heads-up no-limit poker: Libratus beats top professionals
Noam Brown and Tuomas Sandholm · 2017
Cited alongside, same era.
A unified game-theoretic approach to multiagent reinforcement learning
Marc Lanctot, Vinicius Zambaldi, Audrunas Gruslys, Angeliki Lazaridou, Karl Tuyls, Julien Perolat, David Silver, and Thore Graepel · 2017
Cited alongside, same era.
DeepStack: Expert-level artificial intelligence in heads-up no-limit poker
Matej Moravčík and Martin Schmid et al · 2017
Cited alongside, same era.
Deep counterfactual regret minimization
Noam Brown, Adam Lerer, Sam Gross, and Tuomas Sandholm · 2018
Cited alongside, same era.
Solving imperfect-information games via discounted regret minimization
Noam Brown and Tuomas Sandholm · 2018
Cited alongside, same era.
Regret minimization for partially observable deep reinforcement learning
Peter H. Jin, Sergey Levine, and Kurt Keutzer · 2018
Cited alongside, same era.
Double neural counterfactual regret minimization
Hui Li, Kailiang Hu, Zhibang Ge, Tao Jiang, Yuan Qi, and Le Song · 2018
Cited alongside, same era.
Vojtech Kovarík, Martin Schmid, Neil Burch, Michael Bowling, and Viliam Lisý · 2019
Later among the works it cites.
Computing approximate equilibria in sequential adversarial games by exploitability descent
Edward Lockhart, Marc Lanctot, Julien Pérolat, Jean-Baptiste Lespiau, Dustin Morrill, Finbarr Timbers, and Karl Tuyls · 2019
Later among the works it cites.
Variance reduction in monte carlo counterfactual regret minimization (VR-MCCFR) for extensive form games using baselines
Martin Schmid, Neil Burch, Marc Lanctot, Matej Moravcik, Rudolf Kadlec, and Michael Bowling · 2019
Later among the works it cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M. Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H. Choi, Richard Powell, and et al · 2019
Later among the works it cites.
Never give up: Learning directed exploration strategies
Adriá Puigdoménech Badia, Pablo Sprechmann, and et al · 2020
Closest in time.
No-regret learning dynamics for extensive-form correlated and coarse correlated equilibria, 2020
Andrea Celli, Alberto Marchesi, Gabriele Farina, and Nicola Gatti · 2020
Closest in time.
Combining no-regret and q-learning
Ian A Kash, Michael Sullins, and Katja Hofmann · 2020
Closest in time.
Dream: Deep regret minimization with advantage baselines and model-free learning, 2020
Eric Steinberger, Adam Lerer, and Noam Brown · 2020
Closest in time.