Fetching the paper…
Reading the bibliography…
We consider a general class of non-linear Bellman equations.
Dynamic Programming
R. Bellman · 1957
Earlier work this paper cites.
Dynamic programming and Markov processes
R. A. Howard · 1960
Earlier work this paper cites.
Preference reversal and delayed reinforcement
G. Ainslie and R. J. Herrnstein · 1981
Earlier work this paper cites.
Some empirical evidence on dynamic inconsistency
R. Thaler · 1981
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
R. S. Sutton · 1988
Earlier work this paper cites.
Learning from delayed rewards
C. J. C. H. Watkins · 1989
Earlier work this paper cites.
Temporal discounting and preference reversals in choice between delayed outcomes
L. Green, N. Fristoe, and J. Myerson · 1994
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
M. L. Puterman · 1994
Earlier work this paper cites.
TD models: Modeling the world at a mixture of time scales
R. S. Sutton · 1995
Earlier work this paper cites.
Neuro-dynamic Programming
D. P. Bertsekas and J. N. Tsitsiklis · 1996
Cited alongside, same era.
Temporal-difference reinforcement learning with distributed representations
Z. Kurth-Nelson and A. D. Redish · 2009
Cited alongside, same era.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
R. S. Sutton, H. R. Maei, D. Precup, S. Bhatnagar, D. Silver, C. Szepesvári, and E. Wiewiora · 2009
Cited alongside, same era.
Hyperbolically discounted temporal difference learning
W. H. Alexander and J. W. Brown · 2010
Cited alongside, same era.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
R. S. Sutton, J. Modayil, M. Delp, T. Degris, P. M. Pilarski, A. White, and D. Precup · 2011
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
A distributional perspective on reinforcement learning
M. G. Bellemare, W. Dabney, and R. Munos · 2017
Later among the works it cites.
Distributional reinforcement learning with quantile regression
W. Dabney, M. Rowland, M. G. Bellemare, and R. Munos · 2018
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
M. Hessel, J. Modayil, H. Van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. Azar, and D. Silver · 2018
Later among the works it cites.
Observe and look further: Achieving consistent performance on atari
T. Pohlen, B. Piot, T. Hester, M. G. Azar, D. Horgan, D. Budden, G. Barth-Maron, H. van Hasselt, J. Quan, M. Vecerík, M. Hessel, R. Munos, and O. Pietquin · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Cited alongside, same era.
Learning to predict independent of span
H. van Hasselt and R. S. Sutton · 2015
Cited alongside, same era.
Dueling network architectures for deep reinforcement learning
Z. Wang, N. de Freitas, T. Schaul, M. Hessel, H. van Hasselt, and M. Lanctot · 2016
Cited alongside, same era.
Learning values across many orders of magnitude
H. van Hasselt, A. Guez, M. Hessel, V. Mnih, and D. Silver
Cited in the paper.
Deep reinforcement learning with Double Q-learning
H. van Hasselt, A. Guez, and D. Silver
Cited in the paper.
Z. Xu, H. P. van Hasselt, and D. Silver · 2018
Later among the works it cites.
Hyperbolic discounting and learning over multiple horizons
W. Fedus, C. Gelada, Y. Bengio, M. G. Bellemare, and H. Larochelle · 2019
Closest in time.
Recurrent experience replay in distributed reinforcement learning
S. Kapturowski, G. Ostrovski, J. Quan, R. Munos, and W. Dabney · 2019
Closest in time.
Statistics and samples in distributional reinforcement learning
M. Rowland, R. Dadashi, S. Kumar, R. Munos, M. G. Bellemare, and W. Dabney · 2019
Closest in time.