Fetching the paper…
Reading the bibliography…
The popular Q-learning algorithm is known to overestimate action values under certain conditions.
Neocognitron: A hierarchical neural network capable of visual pattern recognition
K. Fukushima · 1988
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
R. S. Sutton · 1988
Earlier work this paper cites.
Learning from delayed rewards
C. J. C. H. Watkins · 1989
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
R. S. Sutton · 1990
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
L. Lin · 1992
Earlier work this paper cites.
Issues in using function approximation for reinforcement learning
S. Thrun and A. Schwartz · 1993
Earlier work this paper cites.
Sample mean based index policies with O(log n) regret for the multi-armed bandit problem
R. Agrawal · 1995
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
L. Baird · 1995
Earlier work this paper cites.
Temporal difference learning and td-gammon
G. Tesauro · 1995
Earlier work this paper cites.
Reinforcement learning: A survey
L. P. Kaelbling, M. L. Littman, and A. W. Moore · 1996
Cited alongside, same era.
An analysis of temporal-difference learning with function approximation
J. N. Tsitsiklis and B. Van Roy · 1997
Cited alongside, same era.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Cited alongside, same era.
Introduction to reinforcement learning
R. S. Sutton and A. G. Barto · 1998
Cited alongside, same era.
Finite-time analysis of the multiarmed bandit problem
P. Auer, N. Cesa-Bianchi, and P. Fischer · 2002
Cited alongside, same era.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
R. I. Brafman and M. Tennenholtz · 2003
Cited alongside, same era.
The many faces of optimism: a unifying approach
I. Szita and A. Lőrincz · 2008
Later among the works it cites.
Reinforcement learning in finite MDPs: PAC analysis
A. L. Strehl, L. Li, and M. L. Littman · 2009
Later among the works it cites.
Double Q-learning
H. van Hasselt · 2010
Later among the works it cites.
Gradient temporal-difference learning algorithms
H. R. Maei · 2011
Later among the works it cites.
Insights in Reinforcement Learning
H. van Hasselt · 2011
Later among the works it cites.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Later among the works it cites.
Human-level control through deep reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforcement learning with factored states and actions
B. Sallans and G. E. Hinton · 2004
Cited alongside, same era.
Neural fitted Q iteration - first experiences with a data efficient neural reinforcement learning method
M. Riedmiller · 2005
Cited alongside, same era.
A convergent O(n) algorithm for off-policy temporal-difference learning with linear function approximation
R. S. Sutton, C. Szepesvári, and H. R. Maei · 2008
Cited alongside, same era.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Closest in time.
Massively parallel methods for deep reinforcement learning
A. Nair, P. Srinivasan, S. Blackwell, C. Alcicek, R. Fearon, A. D. Maria, V. Panneershelvam, M. Suleyman, C. Beattie, S. Petersen, S. Legg, V. Mnih, K. Kavukcuoglu, and D. Silver · 2015
Closest in time.
An emphatic approach to the problem of off-policy temporal-difference learning
R. S. Sutton, A. R. Mahmood, and M. White · 2015
Closest in time.