Fetching the paper…
Reading the bibliography…
Model-free Reinforcement Learning (RL) algorithms such as Q-learning [Watkins, Dayan 92] have been widely used in practice and can achieve human level performance in applications such as video games [Mnih et al.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Finite-sample convergence rates for q-learning and indirect algorithms
Michael J Kearns and Satinder P Singh · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
Similarity estimation techniques from rounding algorithms
Moses S Charikar · 2002
Earlier work this paper cites.
Exploration in metric state spaces
Sham Kakade, Michael J Kearns, and John Langford · 2003
Earlier work this paper cites.
Building triangulations using ϵ \epsilon -nets
Kenneth L Clarkson · 2006
Earlier work this paper cites.
Online regret bounds for undiscounted continuous reinforcement learning
Ronald Ortner and Daniil Ryabko · 2012
Cited alongside, same era.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Cited alongside, same era.
Improved regret bounds for undiscounted continuous reinforcement learning
Kailasam Lakshmanan, Ronald Ortner, and Daniil Ryabko · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Later among the works it cites.
Model-based reinforcement learning in contextual decision processes
Wen Sun, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2018
Later among the works it cites.
Near-optimal time and sample complexities for solving markov decision processes with a generative model
Aaron Sidford, Mengdi Wang, Xian Wu, Lin Yang, and Yinyu Ye · 2018
Later among the works it cites.
Variance reduced value iteration and faster algorithms for solving markov decision processes
Aaron Sidford, Mengdi Wang, Xian Wu, and Yinyu Ye · 2018
Later among the works it cites.
Stephen Tu and Benjamin Recht · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
# Exploration: A study of count-based exploration for deep reinforcement learning
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, Xi Chen, Yan Duan, John Schulman, Filip DeTurck, and Pieter Abbeel · 2017
Cited alongside, same era.
Introduction to multi-armed bandits
Aleksandrs Slivkins · 2019
Closest in time.