Fetching the paper…
Reading the bibliography…
We study algorithms for average-cost reinforcement learning problems with value function approximation.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Adaptive confidence and adaptive curiosity
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
Active exploration in dynamic environments
Sebastian B Thrun and Knut Möller · 1992
Earlier work this paper cites.
Markov decision processes : Discrete stochastic dynamic programming
M. Puterman · 1994
Earlier work this paper cites.
Temporal differences-based policy iteration and applications in neuro-dynamic programming
Dimitri P Bertsekas and Sergey Ioffe · 1996
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
John N. Tsitsiklis and Benjamin Van Roy · 1997
Earlier work this paper cites.
Average cost temporal-difference learning
John N. Tsitsiklis and Benjamin Van Roy · 1999
Earlier work this paper cites.
PAC model-free reinforcement learning
Alexander L. Strehl, Lihong Li, Eric Wiewiora, John Langford, and Michael L. Littman · 2006
Earlier work this paper cites.
Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
A convergent o ( n ) o(n) algorithm for off-policy temporal-difference learning with linear function approximation
R. S. Sutton, Cs. Szepesvári, and H. R. Maei · 2009
Earlier work this paper cites.
Convergence results for some temporal difference methods based on least squares
Huizhen Yu and Dimitri P Bertsekas · 2009
Earlier work this paper cites.
Toward off-policy learning control with function approximation
H. R. Maei, Cs. Szepesvári, S. Bhatnagar, and R. S. Sutton · 2010
Cited alongside, same era.
Convergence of least squares temporal difference methods under general conditions
Huizhen Yu · 2010
Cited alongside, same era.
Error bounds for approximations from projected linear equations
Huizhen Yu and Dimitri P Bertsekas · 2010
Cited alongside, same era.
Approximate policy iteration: A survey and some new methods
Dimitri P Bertsekas · 2011
Cited alongside, same era.
Finite-sample analysis of least-squares policy iteration
Alessandro Lazaric, Mohammad Ghavamzadeh, and Rémi Munos · 2012
Cited alongside, same era.
Regularized off-policy TD-learning
Bo Liu, Sridhar Mahadevan, and Ji Liu · 2012
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Hado van Hasselt, Arthur Guez, and David Silver · 2015
Later among the works it cites.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Nando de Freitas, and Marc Lanctot · 2015
Later among the works it cites.
Regularized policy iteration with nonparametric function spaces
Amir-massoud Farahmand, Mohammad Ghavamzadeh, Csaba Szepesvári, and Shie Mannor · 2016
Later among the works it cites.
Generalization and exploration via randomized value functions
Ian Osband, Zheng Wen, and Benjamin Van Roy · 2016
Later among the works it cites.
Multi-step reinforcement learning: A unifying algorithm
Kristopher De Asis, J. Fernando Hernandez-Garcia, G. Zacharias Holland, and Richard S. Sutton · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Off-policy learning with eligibility traces: A survey
Matthieu Geist and Bruno Scherrer · 2014
Cited alongside, same era.
Finite-sample analysis of proximal gradient TD algorithms
Bo Liu, Ji Liu, Mohammad Ghavamzadeh, Sridhar Mahadevan, and Marek Petrik · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Scale-free algorithms for online linear optimization
Francesco Orabona and Dávid Pál · 2015
Cited alongside, same era.
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2015
Cited alongside, same era.
POLITEX: Regret bounds for policy iteration using expert prediction
Yasin Abbasi-Yadkori, Peter Bartlett, Kush Bhatia, Nevena Lazić, Csaba Szepesvári, and Gellért Weisz
Cited in the paper.
Later among the works it cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Later among the works it cites.
Distributed prioritized experience replay
Dan Horgan, John Quan, David Budden, Gabriel Barth-Maron, Matteo Hessel, Hado Van Hasselt, and David Silver · 2018
Later among the works it cites.
Is Q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sébastien Bubeck, and Michael I. Jordan · 2018
Later among the works it cites.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, et al · 2018
Later among the works it cites.
Large-scale study of curiosity-driven learning
Yuri Burda, Harri Edwards, Deepak Pathak, Amos Storkey, Trevor Darrell, and Alexei A. Efros · 2019
Closest in time.
Provably efficient maximum entropy exploration
Elad Hazan, Sham M Kakade, Karan Singh, and Abby Van Soest · 2019
Closest in time.