Fetching the paper…
Reading the bibliography…
Modern reinforcement learning (RL) commonly engages practical problems with large state spaces, where function approximation must be deployed to approximate either the value function or the policy.
Stochastic games
L. S. Shapley · 1953
Earlier work this paper cites.
The complexity of two-person zero-sum games in extensive form
D. Koller and N. Megiddo · 1992
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
M. L. Littman · 1994
Earlier work this paper cites.
Friend-or-foe q-learning in general-sum games
M. L. Littman · 2001
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
R. I. Brafman and M. Tennenholtz · 2002
Earlier work this paper cites.
Nash q-learning for general-sum stochastic games
J. Hu and M. P. Wellman · 2003
Earlier work this paper cites.
Finding equilibria in large sequential games of imperfect information
A. Gilpin and T. Sandholm · 2006
Earlier work this paper cites.
Regret minimization in games with incomplete information
M. Zinkevich, M. Johanson, M. Bowling, and C. Piccione · 2007
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
T. Jaksch, R. Ortner, and P. Auer · 2010
Earlier work this paper cites.
Algorithms for reinforcement learning
C. Szepesvári · 2010
Earlier work this paper cites.
Competitive Markov decision processes
J. Filar and K. Vrieze · 2012
Earlier work this paper cites.
Swarm robotics: a review from the swarm engineering perspective
M. Brambilla, E. Ferrante, M. Birattari, and M. Dorigo · 2013
Earlier work this paper cites.
Strategy iteration is strongly polynomial for 2-player turn-based stochastic games with a constant discount factor
T. D. Hansen, P. B. Miltersen, and U. Zwick · 2013
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
D. Russo and B. Van Roy · 2013
Earlier work this paper cites.
Model-based reinforcement learning and the eluder dimension
I. Osband and B. Van Roy · 2014
Earlier work this paper cites.
Sample complexity of episodic fixed-horizon reinforcement learning
C. Dann and E. Brunskill · 2015
Earlier work this paper cites.
Pac reinforcement learning with rich observations
A. Krishnamurthy, A. Agarwal, and J. Langford · 2016
Earlier work this paper cites.
Safe, multi-agent, reinforcement learning for autonomous driving
S. Shalev-Shwartz, S. Shammah, and A. Shashua · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
M. G. Azar, I. Osband, and R. Munos · 2017
Earlier work this paper cites.
Contextual decision processes with low bellman rank are pac-learnable
N. Jiang, A. Krishnamurthy, A. Agarwal, J. Langford, and R. E. Schapire · 2017
Earlier work this paper cites.
Multi-agent actor-critic for mixed cooperative-competitive environments
R. Lowe, Y. Wu, A. Tamar, J. Harb, P. Abbeel, and I. Mordatch · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, et al · 2017
Cited alongside, same era.
Online reinforcement learning in stochastic games
C.-Y. Wei, Y.-T. Hong, and C.-J. Lu · 2017
Cited alongside, same era.
Superhuman ai for heads-up no-limit poker: Libratus beats top professionals
N. Brown and T. Sandholm · 2018
Cited alongside, same era.
Is q-learning provably efficient?
C. Jin, Z. Allen-Zhu, S. Bubeck, and M. I. Jordan · 2018
Cited alongside, same era.
Openai five
OpenAI · 2018
Cited alongside, same era.
Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning
T. Rashid, M. Samvelyan, C. Schroeder, G. Farquhar, J. Foerster, and S. Whiteson · 2018
Provably efficient reinforcement learning with linear function approximation
C. Jin, Z. Yang, Z. Wang, and M. I. Jordan · 2020
Later among the works it cites.
Multi-agent trust region policy optimization
H. Li and H. He · 2020
Later among the works it cites.
A sharp analysis of model-based reinforcement learning with self-play
Q. Liu, T. Yu, Y. Bai, and C. Jin · 2020
Later among the works it cites.
A unifying view of optimism in episodic reinforcement learning
G. Neu and C. Pike-Burke · 2020
Later among the works it cites.
Solving discounted stochastic two-player games with near-optimal time and sample complexity
A. Sidford, M. Wang, L. Yang, and Y. Ye · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Dota 2 with large scale deep reinforcement learning
C. Berner, G. Brockman, B. Chan, V. Cheung, P. Debiak, C. Dennison, D. Farhi, Q. Fischer, S. Hashme, C. Hesse, et al · 2019
Cited alongside, same era.
Superhuman ai for multiplayer poker
N. Brown and T. Sandholm · 2019
Cited alongside, same era.
Provably efficient exploration in policy optimization
Q. Cai, Z. Yang, C. Jin, and Z. Wang · 2019
Cited alongside, same era.
Feature-based q-learning for two-player stochastic games
Z. Jia, L. F. Yang, and M. Wang · 2019
Cited alongside, same era.
Model-based rl in contextual decision processes: Pac bounds and exponential improvements over model-free approaches
W. Sun, N. Jiang, A. Krishnamurthy, A. Agarwal, and J. Langford · 2019
Cited alongside, same era.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, et al · 2019
Cited alongside, same era.
R. Wang, R. Salakhutdinov, and L. F. Yang · 2020
Later among the works it cites.
Linear last-iterate convergence in constrained saddle-point optimization
C.-Y. Wei, C.-W. Lee, M. Zhang, and H. Luo · 2020
Later among the works it cites.
G. Weisz, P. Amortila, and C. Szepesvári · 2020
Later among the works it cites.
Learning zero-sum simultaneous-move markov games using function approximation and correlated equilibrium
Q. Xie, Y. Chen, Z. Wang, and Z. Yang · 2020
Later among the works it cites.
Z. Yang, C. Jin, Z. Wang, M. Wang, and M. I. Jordan · 2020
Later among the works it cites.
Learning near optimal policies with low inherent bellman error
A. Zanette, A. Lazaric, M. Kochenderfer, and E. Brunskill · 2020
Later among the works it cites.
Provably efficient reward-agnostic navigation with linear value iteration
A. Zanette, A. Lazaric, M. J. Kochenderfer, and E. Brunskill · 2020
Later among the works it cites.
Model-based multi-agent rl in zero-sum markov games with near-optimal sample complexity
K. Zhang, S. M. Kakade, T. Başar, and L. F. Yang · 2020
Later among the works it cites.
Almost optimal model-free reinforcement learning via reference-advantage decomposition
Z. Zhang, Y. Zhou, and X. Ji · 2020
Later among the works it cites.
Almost optimal algorithms for two-player markov games with linear function approximation
Z. Chen, D. Zhou, and Q. Gu · 2021
Closest in time.
Bilinear classes: A structural framework for provable generalization in rl
S. S. Du, S. M. Kakade, J. D. Lee, S. Lovett, G. Mahajan, W. Sun, and R. Wang · 2021
Closest in time.
Bellman eluder dimension: New rich classes of rl problems, and sample-efficient algorithms
C. Jin, Q. Liu, and S. Miryoosefi · 2021
Closest in time.
Online learning in unknown markov games
Y. Tian, Y. Wang, T. Yu, and S. Sra · 2021
Closest in time.
C.-Y. Wei, C.-W. Lee, M. Zhang, and H. Luo · 2021
Closest in time.
The surprising effectiveness of mappo in cooperative, multi-agent games
C. Yu, A. Velu, E. Vinitsky, Y. Wang, A. Bayen, and Y. Wu · 2021
Closest in time.
Provably efficient algorithms for multi-objective competitive rl
T. Yu, Y. Tian, J. Zhang, and S. Sra · 2021
Closest in time.