Fetching the paper…
Reading the bibliography…
Multi-step methods such as Retrace($\lambda$) and $n$-step $Q$-learning have become a crucial component of modern deep reinforcement learning agents.
“Learning to predict by the methods of temporal differences”
Richard Sutton · 1988
Earlier work this paper cites.
“Learning from Delayed Rewards”, 1989
Christopher Watkins · 1989
Earlier work this paper cites.
“Problem Solving with Reinforcement Learning”, 1995
G.. Rummery · 1995
Earlier work this paper cites.
“Bias-Variance Error Bounds for Temporal Difference Updates”
Michael. Kearns and Satinder. Singh · 2000
Earlier work this paper cites.
“Eligibility traces for off-policy policy evaluation”
D. Precup, R.. Sutton and S. Singh · 2000
Earlier work this paper cites.
“Lecture 6e—RmsProp: Divide the gradient by a running average of its recent magnitude”, COURSERA: Neural Networks for Machine Learning, 2012
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
“Human-level control through deep reinforcement learning”
Volodymyr Mnih et al · 2015
Cited alongside, same era.
“Safe and Efficient Off-Policy Reinforcement Learning”
Rémi Munos, Tom Stepleton, Anna Harutyunyan and Marc. Bellemare · 2016
Cited alongside, same era.
“Multi-step Reinforcement Learning: A Unifying Algorithm”
Kristopher Asis, J. Hernandez-Garcia, G. Holland and Richard. Sutton · 2017
Cited alongside, same era.
“The Reactor: A Sample-Efficient Actor-Critic Architecture”
Audrunas Gruslys, Mohammad Azar, Marc. Bellemare and Rémi Munos · 2017
Cited alongside, same era.
“Rainbow: Combining Improvements in Deep Reinforcement Learning”
Matteo Hessel et al · 2017
Later among the works it cites.
“IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures”
Lasse Espeholt et al · 2018
Later among the works it cites.
“Distributed Prioritized Experience Replay”
Dan Horgan et al · 2018
Later among the works it cites.
“Reinforcement Learning: An Introduction” Manuscript in preparation, 2018
R.. Sutton and A.. Barto · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…