Fetching the paper…
Reading the bibliography…
We provide non-asymptotic bounds for the well-known temporal difference learning algorithm TD(0) with linear function approximators.
Stochastic approximation
David Ruppert · 1991
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Boris T Polyak and Anatoli B Juditsky · 1992
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
John N Tsitsiklis and Benjamin Van Roy · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Actor-Critic Algorithms
Vijay R Konda · 2002
Earlier work this paper cites.
On Actor-Critic Algorithms
Vijay R Konda and John N Tsitsiklis · 2003
Cited alongside, same era.
Natural actor-critic algorithms
S. Bhatnagar, R. Sutton, M. Ghavamzadeh, and M. Lee · 2009
Cited alongside, same era.
Markov chains and stochastic stability
Sean P Meyn and Richard L Tweedie · 2009
Cited alongside, same era.
Convergence results for some temporal difference methods based on least squares
Huizhen Yu and Dimitri P Bertsekas · 2009
Cited alongside, same era.
Finite-sample analysis of lstd
Alessandro Lazaric, Mohammad Ghavamzadeh, and Rémi Munos · 2010
Cited alongside, same era.
Approximate dynamic programming
Dimitri P Bertsekas · 2011
Later among the works it cites.
Concentration Bounds for Stochastic Approximations
Noufel Frikha and Stéphane Menozzi · 2012
Later among the works it cites.
Transport-entropy inequalities and deviation estimates for stochastic approximation schemes
Max Fathi and Noufel Frikha · 2013
Later among the works it cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…