Fetching the paper…
Reading the bibliography…
We design a simple reinforcement learning (RL) agent that implements an optimistic version of $Q$-learning and establish through regret analysis that this agent can operate with some level of competence in any environment.
Discrete dynamic programming
David Blackwell · 1962
Earlier work this paper cites.
Approximations of dynamic programs, I
Ward Whitt · 1978
Earlier work this paper cites.
Optimal control of service rates in networks of queues
Richard R Weber and Shaler Stidham Jr · 1987
Earlier work this paper cites.
A Lagrangian algorithm for computing the optimal service rates in Jackson queuing networks
Kyung Y Jo · 1989
Earlier work this paper cites.
Monotonic and insensitive optimal policies for control of queues with undiscounted costs
Shaler Stidham Jr and Richard R Weber · 1989
Earlier work this paper cites.
Learning from delayed rewards
Christopher John Cornish Hellaby Watkins · 1989
Earlier work this paper cites.
Q-learning
Christopher J.C.H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
On the convergence of stochastic iterative dynamic programming algorithms
Tommi Jaakkola, Michael I Jordan, and Satinder P Singh · 1994
Earlier work this paper cites.
Asynchronous stochastic approximation and Q-learning
John N Tsitsiklis · 1994
Earlier work this paper cites.
Stable function approximation in dynamic programming
Geoffrey J Gordon · 1995
Earlier work this paper cites.
Instance-based utile distinctions for reinforcement learning with hidden state
R Andrew McCallum · 1995
Earlier work this paper cites.
Feature-based methods for large scale dynamic programming
John N Tsitsiklis and Benjamin Van Roy · 1996
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
Universal artificial intelligence: Sequential decisions based on algorithmic probability
Marcus Hutter · 2004
Earlier work this paper cites.
A cost-shaping linear program for average-cost approximate dynamic programming with performance guarantees
Daniela Pucci de Farias and Benjamin Van Roy · 2006
Earlier work this paper cites.
Performance loss bounds for approximate value iteration with state aggregation
Benjamin Van Roy · 2006
Cited alongside, same era.
Approximate and data-driven dynamic programming for queueing networks
Ciamac C Moallemi, Sunil Kumar, and Benjamin Van Roy · 2008
Cited alongside, same era.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Cited alongside, same era.
Stochastic dynamic programming and the control of queueing systems , volume 504
Linn I Sennott · 2009
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Cited alongside, same era.
REGAL: A regularization based algorithm for reinforcement learning in weakly communicating MDPs
Is Q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Later among the works it cites.
Near-optimal representation learning for hierarchical reinforcement learning
Ofir Nachum, Shixiang Gu, Honglak Lee, and Sergey Levine · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Deep exploration via randomized value functions
Ian Osband, Benjamin Van Roy, Daniel J Russo, Zheng Wen, et al · 2019
Later among the works it cites.
Near optimality of finite memory feedback policies in partially observed Markov decision processes
Ali Devran Kara and Serdar Yuksel · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Peter L Bartlett and Ambuj Tewari · 2012
Cited alongside, same era.
Q-learning for history-based reinforcement learning
Mayank Daswani, Peter Sunehag, and Marcus Hutter · 2013
Cited alongside, same era.
(more) efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Cited alongside, same era.
Feature reinforcement learning: state of the art
Mayank Daswani, Peter Sunehag, Marcus Hutter, et al · 2014
Cited alongside, same era.
The dependence of effective planning horizon on model accuracy
Nan Jiang, Alex Kulesza, Satinder Singh, and Richard Lewis · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Mastering Atari, Go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2020
Later among the works it cites.
V-MPO: On-policy maximum a posteriori policy optimization for discrete and continuous control
H. Francis Song, Abbas Abdolmaleki, Jost Tobias Springenberg, Aidan Clark, Hubert Soyer, Jack W. Rae, Seb Noury, Arun Ahuja, Siqi Liu, Dhruva Tirumala, Nicolas Heess, Dan Belov, Martin Riedmiller, and Matthew M. Botvinick · 2020
Later among the works it cites.
Jayakumar Subramanian, Amit Sinha, Raihan Seraj, and Aditya Mahajan · 2020
Later among the works it cites.
Learning and planning in average-reward Markov decision processes
Yi Wan, Abhishek Naik, and Richard S Sutton · 2020
Later among the works it cites.
Model-free reinforcement learning in infinite-horizon average-reward Markov decision processes
Chen-Yu Wei, Mehdi Jafarnia Jahromi, Haipeng Luo, Hiteshi Sharma, and Rahul Jain · 2020
Later among the works it cites.
Zihan Zhang, Xiangyang Ji, and Simon S Du · 2020
Later among the works it cites.
Online learning for unknown partially observable MDPs
Mehdi Jafarnia-Jahromi, Rahul Jain, and Ashutosh Nayyar · 2021
Closest in time.
Reinforcement learning, bit by bit
Xiuyuan Lu, Benjamin Van Roy, Vikranth Dwaracherla, Morteza Ibrahimi, Ian Osband, and Zheng Wen · 2021
Closest in time.
Queue-learning: A reinforcement learning approach for providing quality of service
Majid Raeis, Ali Tizghadam, and Alberto Leon-Garcia · 2021
Closest in time.