Convergence rate of linear two-time-scale stochastic approximation
Vijay R. Konda and John N. Tsitsiklis · 2004
Cited alongside, same era.
Almost sure convergence of two time-scale stochastic approximation algorithms
Vladislav Tadic · 2004
Cited alongside, same era.
Convergence rate and averaging of nonlinear two-time-scale stochastic approximation algorithms
Abdelkader Mokkadem, Mariane Pelletier, et al · 2006
Cited alongside, same era.
Asymptotic analysis of temporal-difference learning algorithms with constant step-sizes
Vladislav Tadic · 2006
Cited alongside, same era.
Stochastic Approximation: A Dynamical Systems Viewpoint
Vivek S Borkar · 2008
Cited alongside, same era.
Advanced Mathematical Tools for Automatic Control Engineers: Deterministic Techniques
Alexander S. Poznyak · 2008
Cited alongside, same era.
Off-policy learning with eligibility traces: A survey
Matthieu Geist and Bruno Scherrer · 2014
Cited alongside, same era.
Finite-sample analysis of proximal gradient td algorithms
Bo Liu, Ji Liu, Mohammad Ghavamzadeh, Sridhar Mahadevan, and Marek Petrik · 2015
Cited alongside, same era.
Finite sample analyses for TD(0) with function approximation
Gal Dalal, Balázs Szörényi, Gugan Thoppe, and Shie Mannor
Cited in the paper.
Finite sample analysis of two-timescale stochastic approximation with applications to reinforcement learning
Gal Dalal, Gugan Thoppe, Balázs Szörényi, and Shie Mannor
Cited in the paper.
A convergent o ( n ) o(n) temporal-difference algorithm for off-policy learning with linear function approximation
Richard S Sutton, Hamid R Maei, and Csaba Szepesvári
Cited in the paper.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Richard S Sutton, Hamid Reza Maei, Doina Precup, Shalabh Bhatnagar, David Silver, Csaba Szepesvári, and Eric Wiewiora
Cited in the paper.