Fetching the paper…
Reading the bibliography…
Estimating the value function for a fixed policy is a fundamental problem in reinforcement learning.
Note on a method for calculating corrected sums of squares and products
BP Welford · 1962
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Leemon Baird · 1995
Earlier work this paper cites.
Generalization in Reinforcement Learning: Safely Approximating the Value Function
J Boyan and A W Moore · 1995
Earlier work this paper cites.
A General Method for Scaling Up Machine Learning Algorithms and its Application to Clustering
Pedro M Domingos and Geoff Hulten · 2001
Earlier work this paper cites.
An Optimal Algorithm for Monte Carlo Estimation
Paul Dagum, Richard Karp, Michael Luby, and Sheldon Ross · 2006
Earlier work this paper cites.
Tuning Bandit Algorithms in Stochastic Environments
Jean-Yves Audibert, Rémi Munos, and Csaba Szepesvari · 2007
Earlier work this paper cites.
Dyna-style planning with linear function approximation and prioritized sweeping
R Sutton, C Szepesvári, A Geramifard, and Michael Bowling · 2008
Earlier work this paper cites.
Empirical Bernstein stopping
Volodymyr Mnih, Csaba Szepesvari, and Jean-Yves Audibert · 2008
Earlier work this paper cites.
Convergent temporal-difference learning with arbitrary smooth function approximation
HR Maei, C Szepesvári, S Bhatnagar, D Precup, D Silver, and Richard S Sutton · 2009
Cited alongside, same era.
TDgamma: Re-evaluating Complex Backups in Temporal Difference Learning
George Konidaris, Scott Niekum, and Philip S Thomas · 2011
Cited alongside, same era.
Off-policy learning with eligibility traces: a survey
Matthieu Geist and Bruno Scherrer · 2014
Cited alongside, same era.
Policy evaluation with temporal differences: a survey and comparison
Christoph Dann, Gerhard Neumann, and Jan Peters · 2014
Cited alongside, same era.
Natural Temporal Difference Learning
William Dabney and Philip S Thomas · 2014
Cited alongside, same era.
Developing a predictive approach to knowledge
Adam White · 2015
Cited alongside, same era.
Concentration inequalities for Markov chains by Marton couplings and spectral methods
Daniel Paulin · 2015
Later among the works it cites.
Investigating practical, linear temporal difference learning
Adam M White and Martha White · 2016
Later among the works it cites.
An emphatic approach to the problem of off-policy temporal-difference learning
Richard S Sutton, A R Mahmood, and Martha White · 2016
Later among the works it cites.
Incremental Truncated LSTD
Clement Gehring, Yangchen Pan, and Martha White · 2016
Later among the works it cites.
Stochastic Variance Reduction Methods for Policy Evaluation
Simon S Du, Jianshu Chen, Lihong Li, Lin Xiao, and Dengyong Zhou · 2017
Later among the works it cites.
Accelerated Gradient Temporal Difference Learning
Yangchen Pan, Adam White, and Martha White · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mixing Time Estimation in Reversible Markov Chains from a Single Sample Path
Daniel Hsu, Aryeh Kontorovich, and Csaba Szepesvari · 2015
Cited alongside, same era.
Learning Sparse Representations in Reinforcement Learning with Sparse Coding
Lei Le, Raksha Kumaraswamy, and Martha White · 2017
Later among the works it cites.