Fetching the paper…
Reading the bibliography…
Model-based reinforcement learning (RL), which finds an optimal policy using an empirical model, has long been recognized as one of the corner stones of RL.
Model-based reinforcement learning with a generative model is minimax optimal
Alekh Agarwal, Sham Kakade, and Lin F Yang · 1906
Earlier work this paper cites.
Optimality and approximation with policy gradient methods in Markov decision processes
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 1908
Earlier work this paper cites.
Stochastic games
Lloyd S Shapley · 1953
Earlier work this paper cites.
Theory of Games and Economic Behavior
John Von Neumann, Oskar Morgenstern, and Harold William Kuhn · 1953
Earlier work this paper cites.
The ellipsoid method and its consequences in combinatorial optimization
Martin Grötschel, László Lovász, and Alexander Schrijver · 1981
Earlier work this paper cites.
Regularity and stability of equilibrium points of bimatrix games
MJM Jansen · 1981
Earlier work this paper cites.
A new polynomial-time algorithm for linear programming
Narendra Karmarkar · 1984
Earlier work this paper cites.
Games and Decisions: Introduction and Critical Survey
R Duncan Luce and Howard Raiffa · 1989
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman · 1994
Earlier work this paper cites.
A Course in Game Theory
Martin J Osborne and Ariel Rubinstein · 1994
Earlier work this paper cites.
Stochastic and Shortest Path Games: Theory and Algorithms
Stephen David Patek · 1997
Earlier work this paper cites.
Linear programming and extensions , volume 48
George Bernard Dantzig · 1998
Earlier work this paper cites.
Finite-sample convergence rates for Q-learning and indirect algorithms
Michael J Kearns and Satinder P Singh · 1999
Earlier work this paper cites.
Variable resolution discretization for high-accuracy solutions of optimal control problems
Remi Munos and Andrew W Moore · 1999
Earlier work this paper cites.
Friend-or-foe Q-learning in general-sum games
Michael L Littman · 2001
Earlier work this paper cites.
R-MAX-A general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
Correlated Q-learning
Amy Greenwald, Keith Hall, and Roberto Serrano · 2003
Earlier work this paper cites.
Nash Q-learning for general-sum stochastic games
Junling Hu and Michael P Wellman · 2003
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham Machandranath Kakade · 2003
Earlier work this paper cites.
An analytic solution to discrete Bayesian reinforcement learning
Pascal Poupart, Nikos Vlassis, Jesse Hoey, and Kevin Regan · 2006
Earlier work this paper cites.
Finite-Dimensional Variational Inequalities and Complementarity Problems
Francisco Facchinei and Jong-Shi Pang · 2007
Earlier work this paper cites.
A comprehensive survey of multi-agent reinforcement learning
Lucian Busoniu, Robert Babuska, and Bart De Schutter · 2008
Earlier work this paper cites.
Reinforcement learning in finite MDPs: PAC analysis
Alexander L Strehl, Lihong Li, and Michael L Littman · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
On the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar, Rémi Munos, and Bert Kappen · 2012
Earlier work this paper cites.
The Implicit Function Theorem: History, Theory, and Applications
Steven G Krantz and Harold R Parks · 2012
Cited alongside, same era.
PAC bounds for discounted mdps
Tor Lattimore and Marcus Hutter · 2012
Cited alongside, same era.
Minimax PAC bounds on the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar, Rémi Munos, and Hilbert J Kappen · 2013
Cited alongside, same era.
Model-based reinforcement learning and the eluder dimension
Ian Osband and Benjamin Van Roy · 2014
Cited alongside, same era.
Sample complexity of episodic fixed-horizon reinforcement learning
Christoph Dann and Emma Brunskill · 2015
Cited alongside, same era.
Learning equilibria of games via payoff queries
John Fearnley, Martin Gairing, Paul W. Goldberg, and Rahul Savani · 2015
Cited alongside, same era.
Near-optimal time and sample complexities for solving Markov decision processes with a generative model
Aaron Sidford, Mengdi Wang, Xian Wu, Lin Yang, and Yinyu Ye · 2018
Later among the works it cites.
Actor-critic policy optimization in partially observable multiagent environments
Sriram Srinivasan, Marc Lanctot, Vinicius Zambaldi, Julien Pérolat, Karl Tuyls, Rémi Munos, and Michael Bowling · 2018
Later among the works it cites.
A theoretical analysis of deep Q-learning
Jianqing Fan, Zhuoran Yang, Yuchen Xie, and Zhaoran Wang · 2019
Later among the works it cites.
Does knowledge transfer always help to learn a better policy?
Fei Feng, Wotao Yin, and Lin F Yang · 2019
Later among the works it cites.
A theory of regularized Markov decision processes
Matthieu Geist, Bruno Scherrer, and Olivier Pietquin · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bayesian reinforcement learning: A survey
Mohammad Ghavamzadeh, Shie Mannor, Joelle Pineau, Aviv Tamar, et al · 2015
Cited alongside, same era.
Approximate dynamic programming for two-player zero-sum Markov games
Julien Pérolat, Bruno Scherrer, Bilal Piot, and Olivier Pietquin · 2015
Cited alongside, same era.
Fast convergence of regularized learning in games
Vasilis Syrgkanis, Alekh Agarwal, Haipeng Luo, and Robert E Schapire · 2015
Cited alongside, same era.
Finding approximate Nash equilibria of bimatrix games via payoff queries
John Fearnley and Rahul Savani · 2016
Cited alongside, same era.
Learning in games via reinforcement and regularization
Panayotis Mertikopoulos and William H Sandholm · 2016
Cited alongside, same era.
Safe, multi-agent, reinforcement learning for autonomous driving
Shai Shalev-Shwartz, Shaked Shammah, and Amnon Shashua · 2016
Cited alongside, same era.
Planning in entropy-regularized Markov decision processes and games
Jean-Bastien Grill, Omar Darwiche Domingues, Pierre Ménard, Rémi Munos, and Michal Valko · 2019
Later among the works it cites.
Feature-based Q-learning for two-player stochastic games
Zeyu Jia, Lin F Yang, and Mengdi Wang · 2019
Later among the works it cites.
Interaction matters: A note on non-asymptotic local convergence of generative adversarial networks
Tengyuan Liang and James Stokes · 2019
Later among the works it cites.
Deep reinforcement learning for cyber security
Thanh Thi Nguyen and Vijay Janapa Reddi · 2019
Later among the works it cites.
AlphaStar: Mastering the Real-Time Strategy Game StarCraft II
Oriol Vinyals, Igor Babuschkin, Junyoung Chung, Michael Mathieu, Max Jaderberg, Wojciech M. Czarnecki, Andrew Dudzik, Aja Huang, Petko Georgiev, Richard Powell, Timo Ewalds, Dan Horgan, Manuel Kroiss, Ivo Danihelka, John Agapiou, Junhyuk Oh, Valentin Dalibard, David Choi, Laurent Sifre, Yury Sulsky, Sasha Vezhnevets, James Molloy, Trevor Cai, David Budden, Tom Paine, Caglar Gulcehre, Ziyu Wang, Tobias Pfaff, Toby Pohlen, Yuhuai Wu, Dani Yogatama, Julia Cohen, Katrina McKinney, Oliver Smith, Tom Schaul, Timothy Lillicrap, Chris Apps, Koray Kavukcuoglu, Demis Hassabis, and David Silver · 2019
Later among the works it cites.
Provable self-play algorithms for competitive reinforcement learning
Yu Bai and Chi Jin · 2020
Closest in time.
Near-optimal reinforcement learning with self-play
Yu Bai, Chi Jin, and Tiancheng Yu · 2020
Closest in time.
Reward-free exploration for reinforcement learning
Chi Jin, Akshay Krishnamurthy, Max Simchowitz, and Tiancheng Yu · 2020
Closest in time.
Breaking the sample size barrier in model-based reinforcement learning with a generative model
Gen Li, Yuting Wei, Yuejie Chi, Yuantao Gu, and Yuxin Chen · 2020
Closest in time.
Deep reinforcement learning for multiagent systems: A review of challenges, solutions, and applications
Thanh Thi Nguyen, Ngoc Duy Nguyen, and Saeid Nahavandi · 2020
Closest in time.
On reinforcement learning for turn-based zero-sum Markov games
Devavrat Shah, Varun Somani, Qiaomin Xie, and Zhi Xu · 2020
Closest in time.
Solving discounted stochastic two-player games with near-optimal time and sample complexity
Aaron Sidford, Mengdi Wang, Lin F Yang, and Yinyu Ye · 2020
Closest in time.
Qiaomin Xie, Yudong Chen, Zhaoran Wang, and Zhuoran Yang · 2020
Closest in time.
V-learning – A simple, efficient, decentralized algorithm for multiagent rl
Chi Jin, Qinghua Liu, Yuanhao Wang, and Tiancheng Yu · 2021
Closest in time.
Global convergence of multi-agent policy gradient in Markov potential games
Stefanos Leonardos, Will Overman, Ioannis Panageas, and Georgios Piliouras · 2021
Closest in time.
A sharp analysis of model-based reinforcement learning with self-play
Qinghua Liu, Tiancheng Yu, Yu Bai, and Chi Jin · 2021
Closest in time.
Independent policy gradient for large-scale Markov potential games: Sharper rates, function approximation, and game-agnostic convergence
Dongsheng Ding, Chen-Yu Wei, Kaiqing Zhang, and Mihailo Jovanovic · 2022
Closest in time.
On improving model-free algorithms for decentralized multi-agent reinforcement learning
Weichao Mao, Lin Yang, Kaiqing Zhang, and Tamer Basar · 2022
Closest in time.
Fictitious play in Markov games with single controller
Muhammed O Sayin, Kaiqing Zhang, and Asuman Ozdaglar · 2022
Closest in time.
Provably efficient reinforcement learning in decentralized general-sum Markov games
Weichao Mao and Tamer Başar · 2023
Closest in time.