Fetching the paper…
Reading the bibliography…
Deep reinforcement learning includes a broad family of algorithms that parameterise an internal representation, such as a value function or policy, by a deep neural network.
Discounted mdp’s: distribution functions and exponential utility maximization
K.-J. Chung and M. J. Sobel · 1987
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
R. S. Sutton · 1988
Earlier work this paper cites.
Adapting bias by gradient descent: An incremental version of delta-bar-delta
R. S. Sutton · 1992
Earlier work this paper cites.
Q-learning
C. J. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
M. L. Puterman · 1994
Earlier work this paper cites.
On-line Q-learning using connectionist systems
G. A. Rummery and M. Niranjan · 1994
Earlier work this paper cites.
Temporal difference learning and TD-Gammon
G. Tesauro · 1995
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Local gain adaptation in stochastic gradient descent
N. N. Schraudolph · 1999
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
G. E. Hinton, S. Osindero, and Y.-W. Teh · 2006
Earlier work this paper cites.
A theoretical and empirical analysis of Expected Sarsa
H. van Seijen, H. van Hasselt, S. Whiteson, and M. Wiering · 2009
Earlier work this paper cites.
Double Q-learning
H. van Hasselt · 2010
Earlier work this paper cites.
Horde: a scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
R. S. Sutton, J. Modayil, M. Delp, T. Degris, P. M. Pilarski, A. White, and D. Precup · 2011
Earlier work this paper cites.
Off-policy actor-critic
T. Degris, M. White, and R. S. Sutton · 2012
Earlier work this paper cites.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
The Arcade Learning Environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Cited alongside, same era.
Recurrent models of visual attention
V. Mnih, N. Heess, A. Graves, and K. Kavukcuoglu · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Cited alongside, same era.
Learning to learn by gradient descent by gradient descent
M. Andrychowicz, M. Denil, S. Gomez, M. W. Hoffman, D. Pfau, T. Schaul, B. Shillingford, and N. De Freitas · 2016
Cited alongside, same era.
Rl2: Fast reinforcement learning via slow reinforcement learning
Y. Duan, J. Schulman, X. Chen, P. L. Bartlett, I. Sutskever, and P. Abbeel · 2016
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
M. Hessel, J. Modayil, H. Van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. Azar, and D. Silver · 2018
Later among the works it cites.
Evolved policy gradients
R. Houthooft, Y. Chen, P. Isola, B. Stadie, F. Wolski, O. J. Ho, and P. Abbeel · 2018
Later among the works it cites.
On first-order meta-learning algorithms
A. Nichol, J. Achiam, and J. Schulman · 2018
Later among the works it cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
Meta-gradient reinforcement learning
Z. Xu, H. P. van Hasselt, and D. Silver · 2018
Later among the works it cites.
On learning intrinsic rewards for policy gradient methods
Z. Zheng, J. Oh, and S. Singh · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforcement learning with unsupervised auxiliary tasks
M. Jaderberg, V. Mnih, W. M. Czarnecki, T. Schaul, J. Z. Leibo, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Deep reinforcement learning with double Q-learning
H. van Hasselt, A. Guez, and D. Silver · 2016
Cited alongside, same era.
Learning to reinforcement learn
J. X. Wang, Z. Kurth-Nelson, D. Tirumala, H. Soyer, J. Z. Leibo, R. Munos, C. Blundell, D. Kumaran, and M. Botvinick · 2016
Cited alongside, same era.
A distributional perspective on reinforcement learning
M. G. Bellemare, W. Dabney, and R. Munos · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
C. Finn, P. Abbeel, and S. Levine · 2017
Cited alongside, same era.
In-Datacenter performance analysis of a Tensor Processing Unit
N. P. Jouppi, C. Young, N. Patil, D. A. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers, R. Boyle, et al · 2017
Cited alongside, same era.
Later among the works it cites.
Meta-learning via learned loss
S. Bechtle, A. Molchanov, Y. Chebotar, E. Grefenstette, L. Righetti, G. Sukhatme, and F. Meier · 2019
Later among the works it cites.
Discovery of useful questions as auxiliary tasks
V. Veeriah, M. Hessel, Z. Xu, J. Rajendran, R. L. Lewis, J. Oh, H. P. van Hasselt, D. Silver, and S. Singh · 2019
Later among the works it cites.
Beyond exponentially discounted sum: Automatic learning of return function
Y. Wang, Q. Ye, and T.-Y. Liu · 2019
Later among the works it cites.
RLax: Reinforcement Learning in JAX, 2020
D. Budden, M. Hessel, J. Quan, and S. Kapturowski · 2020
Closest in time.
Haiku: Sonnet for JAX, 2020
T. Hennigan, T. Cai, T. Norman, and I. Babuschkin · 2020
Closest in time.
Improving generalization in meta reinforcement learning using learned objectives
L. Kirsch, S. van Steenkiste, and J. Schmidhuber · 2020
Closest in time.
Discovering reinforcement learning algorithms
J. Oh, M. Hessel, W. Czarnecki, Z. Xu, H. van Hasselt, S. Singh, and D. Silver · 2020
Closest in time.
Self-tuning deep reinforcement learning
T. Zahavy, Z. Xu, V. Veeriah, M. Hessel, J. Oh, H. van Hasselt, D. Silver, and S. Singh · 2020
Closest in time.
What can learned intrinsic rewards capture?
Z. Zheng, J. Oh, M. Hessel, Z. Xu, M. Kroiss, H. van Hasselt, D. Silver, and S. Singh · 2020
Closest in time.