Fetching the paper…
Reading the bibliography…
One of the main obstacles to broad application of reinforcement learning methods is the parameter sensitivity of our core learning algorithms.
Some studies in machine learning using the game of checkers
A. L. Samuel · 1959
Earlier work this paper cites.
The variance of discounted Markov Decision Processes
M. J. Sobel · 1982
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
R. Sutton · 1988
Earlier work this paper cites.
Self-Improving Reactive Agents Based On Reinforcement Learning, Planning and Teaching
L.-J. Lin · 1992
Earlier work this paper cites.
Practical issues in temporal difference learning
G. Tesauro · 1992
Earlier work this paper cites.
On step-size and bias in temporal-difference learning
R. Sutton and S. P. Singh · 1994
Earlier work this paper cites.
TD models: Modeling the world at a mixture of time scales
R. Sutton · 1995
Earlier work this paper cites.
On the worst-case analysis of temporal-difference learning algorithms
R. E. Schapire and M. K. Warmuth · 1996
Earlier work this paper cites.
Analytical mean squared error curves in temporal difference learning
S. P. Singh and P. Dayan · 1996
Earlier work this paper cites.
Reinforcement learning with replacing eligibility traces
S. P. Singh and R. S. Sutton · 1996
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
J. N. Tsitsiklis and B. Van Roy · 1997
Earlier work this paper cites.
Introduction to reinforcement learning
R. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
R. S. Sutton, D. Precup, and S. Singh · 1999
Cited alongside, same era.
Bias-Variance error bounds for temporal difference updates
M. J. Kearns and S. P. Singh · 2000
Cited alongside, same era.
From Q(lambda) to Average Q-learning: Efficient Implementation of an Asymptotic Approximation
F. Garcia and F. Serre · 2001
Cited alongside, same era.
Bias and variance in value function estimation
S. Mannor, D. Simester, P. Sun, and J. N. Tsitsiklis · 2004
Cited alongside, same era.
A worst-case comparison between temporal difference and residual gradient with linear function approximation
L. Li · 2008
Cited alongside, same era.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Mean-Variance pptimization in Markov Decision Processes
S. Mannor and J. N. Tsitsiklis · 2011
Later among the works it cites.
Adaptive step-size for online temporal difference learning
W. Dabney and A. G. Barto · 2012
Later among the works it cites.
Actor-Critic algorithms for risk-sensitive MDPs
P. L. A and M. Ghavamzadeh · 2013
Later among the works it cites.
Temporal difference methods for the variance of the reward to go
A. Tamar, D. Di Castro, and S. Mannor · 2013
Later among the works it cites.
Weighted importance sampling for off-policy learning with linear function approximation
A. R. Mahmood, H. P. van Hasselt, and R. S. Sutton · 2014
Later among the works it cites.
Multi-timescale nexting in a reinforcement learning robot
J. Modayil, A. White, and R. S. Sutton · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Sutton, H. Maei, D. Precup, and S. Bhatnagar · 2009
Cited alongside, same era.
Temporal difference Bayesian model averaging: A Bayesian perspective on adapting lambda
C. Downey and S. Sanner · 2010
Cited alongside, same era.
GQ ( λ \lambda ): A general gradient algorithm for temporal-difference prediction learning with eligibility traces
H. Maei and R. Sutton · 2010
Cited alongside, same era.
Parametric return density estimation for reinforcement learning
T. Morimura, M. Sugiyama, H. Kashima, H. Hachiya, and T. Tanaka · 2010
Cited alongside, same era.
Interval estimation for reinforcement-learning algorithms in continuous-state domains
M. White and A. White · 2010
Cited alongside, same era.
TDgamma: Re-evaluating Complex Backups in Temporal Difference Learning
G. Konidaris, S. Niekum, and P. S. Thomas · 2011
Cited alongside, same era.
Gradient Temporal-Difference Learning Algorithms
H. Maei · 2011
Cited alongside, same era.
Off-policy TD ( λ \lambda ) with a true online equivalence
H. van Hasselt, A. R. Mahmood, and R. Sutton · 2014
Later among the works it cites.
True online TD(lambda)
H. van Seijen and R. Sutton · 2014
Later among the works it cites.
An emphatic approach to the problem of off-policy temporal-difference learning
R. S. Sutton, A. R. Mahmood, and M. White · 2015
Later among the works it cites.
Developing a predictive approach to knowledge
A. White · 2015
Later among the works it cites.
Transition-based discounting in Markov decision processes
M. White · 2016
Closest in time.