Fetching the paper…
Reading the bibliography…
Two-timescale Stochastic Approximation (SA) algorithms are widely used in Reinforcement Learning (RL).
A projected stochastic approximation method for adaptive filters and identifiers
H Kushner · 1980
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Multivariate stochastic approximation using a simultaneous perturbation gradient approximation
James C Spall · 1992
Earlier work this paper cites.
Rate of convergence of moments of spall’s spsa method
László Gerencsér · 1997
Earlier work this paper cites.
Stochastic Approximation Algorithms and Applications
Harold J. Kushner and G. George Yin · 1997
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
John N Tsitsiklis, Benjamin Van Roy, et al · 1997
Earlier work this paper cites.
Method of variation of parameters for dynamic systems
Vangipuram Lakshmikantham and Sadashiv G Deo · 1998
Earlier work this paper cites.
The ode method for convergence of stochastic approximation and reinforcement learning
Vivek S Borkar and Sean P Meyn · 2000
Earlier work this paper cites.
Actor-Critic Algorithms
Vijaymohan Konda · 2002
Earlier work this paper cites.
Convergence rate of linear two-time-scale stochastic approximation
Vijay R Konda and John N Tsitsiklis · 2004
Cited alongside, same era.
Ordinary differential equations and dynamical systems
Gerald Teschl · 2004
Cited alongside, same era.
Convergence rate and averaging of nonlinear two-time-scale stochastic approximation algorithms
Abdelkader Mokkadem and Mariane Pelletier · 2006
Cited alongside, same era.
Stochastic approximation: a dynamical systems viewpoint
Vivek S Borkar · 2008
Cited alongside, same era.
Natural actor-critic
Jan Peters and Stefan Schaal · 2008
Cited alongside, same era.
Toward off-policy learning control with function approximation
Hamid Reza Maei, Csaba Szepesvári, Shalabh Bhatnagar, and Richard S Sutton · 2010
Cited alongside, same era.
Differential equations, dynamical systems, and an introduction to chaos
Morris W Hirsch, Stephen Smale, and Robert L Devaney · 2012
Later among the works it cites.
Constant step size least-mean-square: Bias-variance trade-offs and optimal sampling distributions
Alexandre Défossez and Francis Bach · 2014
Later among the works it cites.
On td (0) with function approximation: Concentration bounds and a centered variant with exponential convergence
Nathaniel Korda and LA Prashanth · 2015
Later among the works it cites.
Finite-sample analysis of proximal gradient td algorithms
Bo Liu, Ji Liu, Mohammad Ghavamzadeh, Sridhar Mahadevan, and Marek Petrik · 2015
Later among the works it cites.
An emphatic approach to the problem of off-policy temporal-difference learning
Richard S Sutton, A Rupam Mahmood, and Martha White · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The fixed points of off-policy td
J Zico Kolter · 2011
Cited alongside, same era.
Gradient temporal-difference learning algorithms
Hamid Reza Maei · 2011
Cited alongside, same era.
Dynamic Programming and Optimal Control
D. P. Bertsekas · 2012
Cited alongside, same era.
Convergent temporal-difference learning with arbitrary smooth function approximation
Shalabh Bhatnagar, Doina Precup, David Silver, Richard S Sutton, Hamid R Maei, and Csaba Szepesvári
Cited in the paper.
Natural actor-critic algorithms
Shalabh Bhatnagar, Richard Sutton, Mohammad Ghavamzadeh, and Mark Lee
Cited in the paper.
A convergent o(n) temporal-difference algorithm for off-policy learning with linear function approximation
Richard S Sutton, Hamid R Maei, and Csaba Szepesvári
Cited in the paper.
Gugan Thoppe and Vivek S Borkar · 2015
Later among the works it cites.
A stability criterion for two timescale stochastic approximation schemes
Chandrashekar Lakshminarayanan and Shalabh Bhatnagar · 2017
Closest in time.
Finite sample analyses for td(0) with function approximation
Gal Dalal, Balazs Szorenyi, Gugan Thoppe, and Shie Mannor · 2018
Closest in time.
Linear stochastic approximation: How far does constant step-size and iterate averaging go?
Chandrashekar Lakshminarayanan and Csaba Szepesvari · 2018
Closest in time.