Fetching the paper…
Reading the bibliography…
We consider how to learn multi-step predictions efficiently.
“A stochastic approximation method”
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
“Dynamic Programming”
R. Bellman · 1957
Earlier work this paper cites.
“Dynamic programming and Markov processes”
R.˜A. Howard · 1960
Earlier work this paper cites.
“Temporal credit assignment in reinforcement learning”, 1984
R.˜S. Sutton · 1984
Earlier work this paper cites.
“Learning to predict by the methods of temporal differences”
R.˜S. Sutton · 1988
Earlier work this paper cites.
“Simultaneous localization and mapping with unknown data association using FastSLAM”
Michael Montemerlo and Sebastian Thrun · 1991
Earlier work this paper cites.
“Acceleration of stochastic approximation by averaging”
Boris˜T Polyak and Anatoli˜B Juditsky · 1992
Earlier work this paper cites.
“Linear least-squares algorithms for temporal difference learning”
S.˜J. Bradtke and A.˜G. Barto · 1996
Earlier work this paper cites.
“Stochastic approximation with two time scales”
Vivek˜S Borkar · 1997
Earlier work this paper cites.
“Reinforcement Learning: An Introduction”
R.˜S. Sutton and A.˜G. Barto · 1998
Cited alongside, same era.
“Eligibility traces for off-policy policy evaluation”
D. Precup, R.˜S. Sutton and S.˜P. Singh · 2000
Cited alongside, same era.
“Off-policy temporal-difference learning with function approximation”
D. Precup and R.˜S. Sutton · 2001
Cited alongside, same era.
“Stochastic approximation and recursive algorithms and applications”
Harold˜J Kushner and George Yin · 2003
Cited alongside, same era.
“Convergence rate of linear two-time-scale stochastic approximation”
Vijay˜R Konda and John˜N Tsitsiklis · 2004
Cited alongside, same era.
“Stochastic approximation”
Vivek˜S Borkar · 2008
Cited alongside, same era.
“Gradient temporal-difference learning algorithms”, 2011
Hamid˜Reza Maei · 2011
Later among the works it cites.
“Non-strongly-convex smooth stochastic approximation with convergence rate O ( 1 / n ) O(1/n) ”
Francis Bach and Eric Moulines · 2013
Later among the works it cites.
“Weighted importance sampling for off-policy learning with linear function approximation”
A˜Rupam Mahmood, Hado˜P van Hasselt and Richard˜S Sutton · 2014
Later among the works it cites.
“A new Q( λ \lambda ) with interim forward view and Monte Carlo equivalence”
Rich˜S Sutton, Ashique˜R Mahmood, Doina Precup and Hado˜P van Hasselt · 2014
Later among the works it cites.
“Off-policy TD( λ \lambda ) with a true online equivalence”
Hado˜P van Hasselt, A˜Rupam Mahmood and Richard˜S Sutton · 2014
Later among the works it cites.
“True online TD( λ \lambda )”
Harm van Seijen and Rich˜S Sutton · 2014
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“A convergent O(n) algorithm for off-policy temporal-difference learning with linear function approximation”
R.˜S. Sutton, . Szepesv“’ari and H.˜R. Maei · 2008
Cited alongside, same era.
“Fast gradient-descent methods for temporal-difference learning with linear function approximation”
R.˜S. Sutton, H.˜R. Maei, D. Precup, S. Bhatnagar, D. Silver, . Szepesv“’ari and E. Wiewiora · 2009
Cited alongside, same era.
“Algorithms for reinforcement learning”
aba Szepesv“’ari · 2010
Cited alongside, same era.
Later among the works it cites.
“Deep learning”
Yann LeCun, Yoshua Bengio and Geoffrey Hinton · 2015
Closest in time.
“Human-level control through deep reinforcement learning”
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei˜A. Rusu, Joel Veness, Marc˜G. Bellemare, Alex Graves, Martin Riedmiller, Andreas˜K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg and Demis Hassabis · 2015
Closest in time.