Fetching the paper…
Reading the bibliography…
Reinforcement learning tasks are typically specified as Markov decision processes.
TD models: Modeling the world at a mixture of time scales
Richard S Sutton · 1995
Earlier work this paper cites.
Neuro-dynamic programming
Dimitri P Bertsekas and John N Tsitsiklis · 1996
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
Johnathan N Tsitsiklis and Benjamin Van Roy · 1997
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S Sutton and A G Barto · 1998
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Hierarchical Reinforcement Learning with the MAXQ Value Function Decomposition
Thomas G Dietterich · 2000
Earlier work this paper cites.
Bias-Variance error bounds for temporal difference updates
Michael J Kearns and Satinder P Singh · 2000
Earlier work this paper cites.
An object-oriented representation for efficient reinforcement learning
Carlos Diuk, Andre Cohen, and Michael L Littman · 2008
Earlier work this paper cites.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Richard S Sutton, Hamid Maei Reza, Doina Precup, and Shalab Bhatnagar · 2009
Cited alongside, same era.
GQ ( λ \lambda ): A general gradient algorithm for temporal-difference prediction learning with eligibility traces
Hamid Maei Reza and Richard S Sutton · 2010
Cited alongside, same era.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Richard S Sutton, Joseph Modayil, Michael Delp, Thomas Degris, Patrick M Pilarski, Adam White, and Doina Precup · 2011
Cited alongside, same era.
Insights in Reinforcement Learning
Hado Philip van Hasselt · 2011
Cited alongside, same era.
Least Squares Temporal Difference Methods: An Analysis under General Conditions
Huizhen Yu · 2012
Cited alongside, same era.
Generalized Emphatic Temporal Difference Learning: Bias-Variance Analysis
Assaf Hallak, Aviv Tamar, Rémi Munos, and Shie Mannor · 2015
Later among the works it cites.
The Dependence of Effective Planning Horizon on Model Accuracy
Nan Jiang, Alex Kulesza, Satinder P Singh, and Richard L Lewis · 2015
Later among the works it cites.
Emphatic temporal-difference learning
Ashique Rupam Mahmood, Huizhen Yu, Martha White, and Richard S Sutton · 2015
Later among the works it cites.
On the Rate of Convergence and Error Bounds for LSTD( λ \lambda )
Manel Tagorti and Bruno Scherrer · 2015
Later among the works it cites.
Developing a predictive approach to knowledge
Adam White · 2015
Later among the works it cites.
On convergence of emphatic temporal-difference learning
Huizhen Yu · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Joseph Modayil, Adam White, and Richard S Sutton · 2014
Cited alongside, same era.
A new Q(lambda) with interim forward view and Monte Carlo equivalence
Richard S Sutton, Ashique Rupam Mahmood, Doina Precup, and Hado van Hasselt · 2014
Cited alongside, same era.
True online TD(lambda)
Harm van Seijen and Rich Sutton · 2014
Cited alongside, same era.
Later among the works it cites.
An emphatic approach to the problem of off-policy temporal-difference learning
Richard S Sutton, Ashique Rupam Mahmood, and Martha White · 2016
Closest in time.