Fetching the paper…
Reading the bibliography…
Temporal difference learning (TD) is a simple iterative algorithm used to estimate the value function corresponding to a given policy in a Markov decision process.
Stochastic optimal control: the discrete time case
Dimitri P Bertsekas and Steven E Shreve · 1978
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Convergence of stochastic iterative dynamic programming algorithms
Tommi Jaakkola, Michael I Jordan, and Satinder P Singh · 1994
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Leemon Baird · 1995
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Steven J Bradtke and Andrew G Barto · 1996
Earlier work this paper cites.
On the averaged stochastic approximation for linear regression
László Györfi and Harro Walk · 1996
Earlier work this paper cites.
On the worst-case analysis of temporal-difference learning algorithms
Robert E Schapire and Manfred K Warmuth · 1996
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
John N Tsitsiklis and Benjamin Van Roy · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Learning and value function approximation in complex decision processes
Benjamin Van Roy · 1998
Earlier work this paper cites.
Optimal stopping of markov processes: Hilbert space theory, approximation algorithms, and an application to pricing high-dimensional financial derivatives
John N Tsitsiklis and Benjamin Van Roy · 1999
Earlier work this paper cites.
The ODE method for convergence of stochastic approximation and reinforcement learning
Vivek S Borkar and Sean P Meyn · 2000
Earlier work this paper cites.
Actor-Critic Algorithms
Vijay R Konda · 2002
Earlier work this paper cites.
The linear programming approach to approximate dynamic programming
Daniela Pucci De Farias and Benjamin Van Roy · 2003
Earlier work this paper cites.
Stochastic approximation and recursive algorithms and applications , volume 35
Harold Kushner and Gang G Yin · 2003
Earlier work this paper cites.
Primal-dual simulation algorithm for pricing multidimensional american options
Leif Andersen and Mark Broadie · 2004
Earlier work this paper cites.
Pricing american options: A duality approach
Martin B Haugh and Leonid Kogan · 2004
Cited alongside, same era.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Cited alongside, same era.
Stochastic approximation: a dynamical systems viewpoint , volume 48
Vivek S Borkar · 2009
Cited alongside, same era.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Cited alongside, same era.
Convergence results for some temporal difference methods based on least squares
Huizhen Yu and Dimitri P Bertsekas · 2009
Cited alongside, same era.
LSTD with Random Projections
Policy evaluation with temporal differences: A survey and comparison
Christoph Dann, Gerhard Neumann, and Jan Peters · 2014
Later among the works it cites.
True online td (lambda)
Harm Seijen and Richard S Sutton · 2014
Later among the works it cites.
Convex optimization: Algorithms and complexity
Sébastien Bubeck · 2015
Later among the works it cites.
On TD(0) with function approximation: Concentration bounds and a centered variant with exponential convergence
Nathaniel Korda and Prashanth La · 2015
Later among the works it cites.
Finite-sample analysis of proximal gradient td algorithms
Bo Liu, Mohammad Liu, Ji ]and Ghavamzadeh, Sridhar Mahadevan, and Marek Petrik · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mohammad Ghavamzadeh, Alessandro Lazaric, Odalric Maillard, and Rémi Munos · 2010
Cited alongside, same era.
Stochastic approximation: A survey
Harold Kushner · 2010
Cited alongside, same era.
Finite-sample analysis of LSTD
Alessandro Lazaric, Mohammad Ghavamzadeh, and Rémi Munos · 2010
Cited alongside, same era.
Adaptive algorithms and stochastic approximations , volume 22
Albert Benveniste, Michel Métivier, and Pierre Priouret · 2012
Cited alongside, same era.
Dynamic programming and optimal control , volume 2
Dimitri P Bertsekas · 2012
Cited alongside, same era.
Pathwise optimization for optimal stopping problems
Vijay V Desai, Vivek F Farias, and Ciamac C Moallemi · 2012
Cited alongside, same era.
Simon Lacoste-Julien, Mark Schmidt, and Francis Bach · 2012
Cited alongside, same era.
Later among the works it cites.
Zap q-learning
Adithya M Devraj and Sean P Meyn · 2017
Later among the works it cites.
Non-convex optimization for machine learning
Prateek Jain and Purushottam Kar · 2017
Later among the works it cites.
Finite Time Bounds for Temporal Difference Learning with Function Approximation: Problems with some “state-of-the-art” results, 2017
Chandrashekar Lakshminarayanan and Csaba Szepesvári · 2017
Later among the works it cites.
Markov chains and mixing times , volume 107
David A Levin and Yuval Peres · 2017
Later among the works it cites.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Closest in time.
Beating the curse of dimensionality in options pricing and optimal stopping
David A Goldberg and Yilun Chen · 2018
Closest in time.
Linear stochastic approximation: How far does constant step-size and iterate averaging go?
Chandrashekar Lakshminarayanan and Csaba Szepesvári · 2018
Closest in time.
Convergent TREE BACKUP and RETRACE with function approximation
Ahmed Touati, Pierre-Luc Bacon, Doina Precup, and Pascal Vincent · 2018
Closest in time.
Least-squares temporal difference learning for the linear quadratic regulator
Stephen Tu and Benjamin Recht · 2018
Closest in time.