Fetching the paper…
Reading the bibliography…
We consider $d$-dimensional linear stochastic approximation algorithms (LSAs) with a constant step-size and the so called Polyak-Ruppert (PR) averaging of iterates.
Efficient estimations from a slowly convergent robbins-monro process
David Ruppert · 1988
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Long-Ji Lin · 1992
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Boris T Polyak and Anatoli B Juditsky · 1992
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
On average versus discounted reward temporal-difference learning
John N Tsitsiklis and Benjamin Van Roy · 2002
Earlier work this paper cites.
Linear stochastic approximation driven by slowly varying Markov chains
Vijay R Konda and John N Tsitsiklis · 2003
Cited alongside, same era.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Francis R Bach and Eric Moulines · 2011
Cited alongside, same era.
Batch reinforcement learning
Sascha Lange, Thomas Gabel, and Martin Riedmiller · 2012
Cited alongside, same era.
Policy evaluation with temporal differences: a survey and comparison
Christoph Dann, Gerhard Neumann, and Jan Peters · 2014
Cited alongside, same era.
Averaged least-mean-squares: Bias-variance trade-offs and optimal sampling distributions
Alexandre Défossez and Francis Bach · 2015
Cited alongside, same era.
A convergent o ( n ) o(n) temporal-difference algorithm for off-policy learning with linear function approximation
Richard S Sutton, Hamid R Maei, and Csaba Szepesvári
Cited in the paper.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Richard S Sutton, Hamid Reza Maei, Doina Precup, Shalabh Bhatnagar, David Silver, Csaba Szepesvári, and Eric Wiewiora
Cited in the paper.
On TD(0) with function approximation: Concentration bounds and a centered variant with exponential convergence
Nathaniel Korda and LA Prashanth · 2015
Later among the works it cites.
Finite-sample analysis of proximal gradient td algorithms
Bo Liu, Ji Liu, Mohammad Ghavamzadeh, Sridhar Mahadevan, and Marek Petrik · 2015
Later among the works it cites.
Harder, better, faster, stronger convergence rates for least-squares regression
Aymeric Dieuleveut, Nicolas Flammarion, and Francis Bach · 2016
Later among the works it cites.
State of the art control of atari games using shallow reinforcement learning
Yitao Liang, Marlos C. Machado, Erik Talvitie, and Michael H. Bowling · 2016
Later among the works it cites.
Stochastic variance reduction methods for policy evaluation
Simon S Du, Jianshu Chen, Lihong Li, Lin Xiao, and Dengyong Zhou · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…