Fetching the paper…
Reading the bibliography…
The goal of reinforcement learning algorithms is to estimate and/or optimise the value function.
Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook
J. Schmidhuber · 1987
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
R. S. Sutton · 1988
Earlier work this paper cites.
Adapting bias by gradient descent: An incremental version of delta-bar-delta
R. S. Sutton · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
On-line Q-learning using connectionist sytems
G. A. Rummery and M. Niranjan · 1994
Earlier work this paper cites.
Neuro-Dynamic Programming
D. P. Bertsekas and J. N. Tsitsiklis · 1996
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Adaptive critic designs
D. V. Prokhorov and D. C. Wunsch · 1997
Earlier work this paper cites.
Learning to drive a bicycle using reinforcement learning and shaping
J. Randløv and P. Alstrøm · 1998
Earlier work this paper cites.
Analytical mean squared error curves for temporal difference learning
S. Singh and P. Dayan · 1998
Earlier work this paper cites.
Learning to learn
S. Thrun and L. Pratt · 1998
Earlier work this paper cites.
Local gain adaptation in stochastic gradient descent
N. N. Schraudolph · 1999
Earlier work this paper cites.
Bias-variance error bounds for temporal difference updates
M. J. Kearns and S. P. Singh · 2000
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
D. Precup, R. S. Sutton, and S. P. Singh · 2000
Earlier work this paper cites.
Learning to learn using gradient descent
S. Hochreiter, A. S. Younger, and P. R. Conwell · 2001
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
S. P. Singh, A. G. Barto, and N. Chentanez · 2005
Earlier work this paper cites.
A theoretical and empirical analysis of Expected Sarsa
H. van Seijen, H. van Hasselt, S. Whiteson, and M. Wiering · 2009
Earlier work this paper cites.
Temporal difference bayesian model averaging: A bayesian perspective on adapting lambda
C. Downey and S. Sanner · 2010
Earlier work this paper cites.
Double Q-learning
H. van Hasselt · 2010
Cited alongside, same era.
Practical Bayesian optimization of machine learning algorithms
J. Snoek, H. Larochelle, and R. P. Adams · 2012
Cited alongside, same era.
Lecture 6.5-RMSProp: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Cited alongside, same era.
A new Q( λ \lambda ) with interim forward view and Monte Carlo equivalence
R. S. Sutton, A. R. Mahmood, D. Precup, and H. van Hasselt · 2014
Cited alongside, same era.
Hyperparameter optimization with approximate gradient
F. Pedregosa · 2016
Later among the works it cites.
Prioritized experience replay
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2016
Later among the works it cites.
An emphatic approach to the problem of off-policy temporal-difference learning
R. S. Sutton, A. R. Mahmood, and M. White · 2016
Later among the works it cites.
Deep reinforcement learning with double Q-learning
H. van Hasselt, A. Guez, and D. Silver · 2016
Later among the works it cites.
A greedy approach to adapting the trace parameter for temporal difference learning
M. White and A. White · 2016
Later among the works it cites.
A distributional perspective on reinforcement learning
M. G. Bellemare, W. Dabney, and R. Munos · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
ADAM: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Cited alongside, same era.
Gradient-based hyperparameter optimization through reversible learning
D. Maclaurin, D. Duvenaud, and R. Adams · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
Massively parallel methods for deep reinforcement learning
A. Nair, P. Srinivasan, S. Blackwell, C. Alcicek, R. Fearon, A. De Maria, V. Panneershelvam, M. Suleyman, C. Beattie, S. Petersen, et al · 2015
Cited alongside, same era.
Universal value function approximators
T. Schaul, D. Horgan, K. Gregor, and D. Silver · 2015
Cited alongside, same era.
TensorFlow: A system for large-scale machine learning
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, et al · 2016
Cited alongside, same era.
Online meta-learning by parallel algorithm competition
S. Elfwing, E. Uchibe, and K. Doya · 2017
Later among the works it cites.
Forward and reverse gradient-based hyperparameter optimization
L. Franceschi, M. Donini, P. Frasconi, and M. Pontil · 2017
Later among the works it cites.
Incremental Off-policy Reinforcement Learning Algorithms
A. Mahmood · 2017
Later among the works it cites.
Learned optimizers that scale and generalize
O. Wichrowska, N. Maheswaranathan, M. W. Hoffman, S. G. Colmenarejo, M. Denil, N. de Freitas, and J. Sohl-Dickstein · 2017
Later among the works it cites.
Continuous adaptation via meta-learning in nonstationary and competitive environments
M. Al-Shedivat, T. Bansal, Y. Burda, I. Sutskever, I. Mordatch, and P. Abbeel · 2018
Closest in time.
IMPALA: Scalable distributed Deep-RL with importance weighted actor-learner architectures
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, et al · 2018
Closest in time.
Meta-learning and universality: Deep representations and gradient descent can approximate any learning algorithm
C. Finn and S. Levine · 2018
Closest in time.
Noisy networks for exploration
M. Fortunato, M. G. Azar, B. Piot, J. Menick, I. Osband, A. Graves, V. Mnih, R. Munos, D. Hassabis, O. Pietquin, et al · 2018
Closest in time.
Recasting gradient-based meta-learning as hierarchical Bayes
E. Grant, C. Finn, S. Levine, T. Darrell, and T. Griffiths · 2018
Closest in time.
Rainbow: Combining improvements in deep reinforcement learning
M. Hessel, J. Modayil, H. van Hasselt, T. Schaul, G. Ostrovski, W. Dabney, D. Horgan, B. Piot, M. Azar, and D. Silver · 2018
Closest in time.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Closest in time.
On learning intrinsic rewards for policy gradient methods
Z. Zheng, J. Oh, and S. Singh · 2018
Closest in time.