Fetching the paper…
Reading the bibliography…
We study stochastic approximation procedures for approximately solving a $d$-dimensional linear fixed point equation based on observing a trajectory of length $n$ from an ergodic Markov chain.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Statistical methods in Markov chains
Patrick Billingsley · 1961
Earlier work this paper cites.
First and second order asymptotic efficiencies of estimators
C. R. Rao · 1962
Earlier work this paper cites.
Analysis of recursive stochastic algorithms
Lennart Ljung · 1977
Earlier work this paper cites.
On positive real transfer functions and the convergence of some recursive schemes
Lennart Ljung · 1977
Earlier work this paper cites.
Stochastic approximation methods for constrained and unconstrained systems
Harold J. Kushner and Dean S. Clark · 1978
Earlier work this paper cites.
Second order efficiency of the mle with respect to any bounded bowl-shape loss function
JK Ghosh, Bimal K Sinha, and HS Wieand · 1980
Earlier work this paper cites.
Applications of a Kushner and Clark lemma to general classes of stochastic algorithms
Michel Metivier and Pierre Priouret · 1984
Earlier work this paper cites.
Efficient estimations from a slowly convergent Robbins-Monro process
David Ruppert · 1988
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
On a problem of adaptive estimation in Gaussian white noise
OV Lepskii · 1991
Earlier work this paper cites.
Efficient estimation of the stationary distribution for exponentially ergodic Markov chains
Spiridon Penev · 1991
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Boris T Polyak and Anatoli B Juditsky · 1992
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
TD( λ \lambda ) converges with probability 1
Peter Dayan and Terrence J Sejnowski · 1994
Earlier work this paper cites.
On-line Q-learning using connectionist systems
Gavin A Rummery and Mahesan Niranjan · 1994
Earlier work this paper cites.
Asynchronous stochastic approximation and Q Q -learning
J. N. Tsitsiklis · 1994
Earlier work this paper cites.
Applications of the van Trees inequality: a Bayesian Cramér-Rao bound
Richard D Gill and Boris Y Levit · 1995
Earlier work this paper cites.
Efficiency of empirical estimators for Markov chains
Priscilla E Greenwood and Wolfgang Wefelmeyer · 1995
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Steven J Bradtke and Andrew G Barto · 1996
Earlier work this paper cites.
Efficiency of the empirical distribution for ergodic diffusion
Yury A Kutoyants · 1997
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
John N Tsitsiklis and Benjamin Van Roy · 1997
Earlier work this paper cites.
The asymptotic convergence-rate of Q-learning
Csaba Szepesvári · 1998
Earlier work this paper cites.
Optimal stopping of Markov processes: Hilbert space theory, approximation algorithms, and an application to pricing high-dimensional financial derivatives
John N Tsitsiklis and Benjamin Van Roy · 1999
Earlier work this paper cites.
Actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
Asymptotic Statistics
Aad W van der Vaart · 2000
Earlier work this paper cites.
Lectures on modern convex optimization
Arkadi Nemirovski · 2001
Earlier work this paper cites.
Technical update: Least-squares temporal difference learning
Justin A Boyan · 2002
Earlier work this paper cites.
Learning rates for Q Q -learning
E. Even-Dar and Y. Mansour · 2003
Earlier work this paper cites.
Stochastic approximation and recursive algorithms and applications
Harold Kushner and G George Yin · 2003
Cited alongside, same era.
New introduction to multiple time series analysis
H. Lütkepohl · 2005
Cited alongside, same era.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Cited alongside, same era.
Introduction to Nonparametric Estimation
Alexandre B Tsybakov · 2008
Cited alongside, same era.
Time series: theory and methods
Peter J Brockwell and Richard A Davis · 2009
Cited alongside, same era.
Stochastic Approximation: a Dynamical Systems Viewpoint
Vivek S Borkar · 2009
Cited alongside, same era.
Robust stochastic approximation approach to stochastic programming
Finite-time error bounds for linear stochastic approximation and TD learning
Rayadurgam Srikant and Lei Ying · 2019
Later among the works it cites.
High-dimensional Statistics: A Non-asymptotic Viewpoint
Martin J Wainwright · 2019
Later among the works it cites.
Martin J Wainwright · 2019
Later among the works it cites.
Variance-reduced Q-learning is minimax optimal
Martin J Wainwright · 2019
Later among the works it cites.
Least squares regression with Markovian data: Fundamental limits and algorithms
Guy Bresler, Prateek Jain, Dheeraj Nagaraj, Praneeth Netrapalli, and Xian Wu · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Cited alongside, same era.
Algorithms for Reinforcement Learning
Csaba Szepesvári · 2010
Cited alongside, same era.
Error bounds for approximations from projected linear equations
Huizhen Yu and Dimitri P Bertsekas · 2010
Cited alongside, same era.
Temporal difference methods for general projected equations
Dimitri P Bertsekas · 2011
Cited alongside, same era.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Éric Moulines and Francis R Bach · 2011
Cited alongside, same era.
Adaptive Algorithms and Stochastic Approximations
Albert Benveniste, Michel Métivier, and Pierre Priouret · 2012
Cited alongside, same era.
Coupling and convergence for Hamiltonian Monte Carlo
Nawaf Bou-Rabee, Anreas Eberle, and Raphael Zimmer · 2020
Later among the works it cites.
Explicit mean-square error bounds for Monte-Carlo and linear stochastic approximation
Shuhang Chen, Adithya Devraj, Ana Busic, and Sean Meyn · 2020
Later among the works it cites.
Statistical inference for model parameters in stochastic gradient descent
Xi Chen, Jason D Lee, Xin T Tong, and Yichen Zhang · 2020
Later among the works it cites.
Bridging the gap between constant step size stochastic gradient descent and Markov chains
Aymeric Dieuleveut, Alain Durmus, and Francis Bach · 2020
Later among the works it cites.
Finite-time analysis of stochastic gradient descent under Markov randomness
Thinh T Doan, Lam M Nguyen, Nhan H Pham, and Justin Romberg · 2020
Later among the works it cites.
Time series analysis
James Douglas Hamilton · 2020
Later among the works it cites.
Georgios Kotsalis, Guanghui Lan, and Tianjiao Li · 2020
Later among the works it cites.
Finite time analysis of linear two-timescale stochastic approximation with Markovian noise
Maxim Kaledin, Eric Moulines, Alexey Naumov, Vladislav Tadic, and Hoi-To Wai · 2020
Later among the works it cites.
ROOT-SGD: Sharp nonasymptotics and asymptotic efficiency in a single algorithm
Chris Junchi Li, Wenlong Mou, Martin J Wainwright, and Michael I Jordan · 2020
Later among the works it cites.
Breaking the sample size barrier in model-based reinforcement learning with a generative model
Gen Li, Yuting Wei, Yuejie Chi, Yuantao Gu, and Yuxin Chen · 2020
Later among the works it cites.
On linear stochastic approximation: Fine-grained Polyak-Ruppert and non-asymptotic concentration
Wenlong Mou, Chris Junchi Li, Martin J Wainwright, Peter L Bartlett, and Michael I Jordan · 2020
Later among the works it cites.
Optimal oracle inequalities for solving projected fixed-point equations
Wenlong Mou, Ashwin Pananjady, and Martin J Wainwright · 2020
Later among the works it cites.
An analysis of constant step size SGD in the non-convex regime: Asymptotic normality and bias
Lu Yu, Krishnakumar Balasubramanian, Stanislav Volgushev, and Murat A Erdogdu · 2020
Later among the works it cites.
A concentration bound for contractive stochastic approximation
Vivek S Borkar · 2021
Closest in time.
A Lyapunov theory for finite-sample guarantees of asynchronous Q-learning and TD-learning variants
Zaiwei Chen, Siva Theja Maguluri, Sanjay Shakkottai, and Karthikeyan Shanmugam · 2021
Closest in time.
On the convergence of stochastic approximations under a subgeometric ergodic Markov dynamic
Vianney Debavelaere, Stanley Durrleman, and Stéphanie Allassonnière · 2021
Closest in time.
Alain Durmus, Eric Moulines, Alexey Naumov, Sergey Samsonov, and Hoi-To Wai · 2021
Closest in time.
Optimal policy evaluation using kernel-based temporal difference methods
Yaqi Duan, Mengdi Wang, and Martin J Wainwright · 2021
Closest in time.
Streaming linear system identification with reverse experience replay
Prateek Jain, Suhas S Kowshik, Dheeraj Nagaraj, and Praneeth Netrapalli · 2021
Closest in time.
Is temporal difference learning optimal? An instance-dependent analysis
Koulik Khamaru, Ashwin Pananjady, Feng Ruan, Martin J Wainwright, and Michael I Jordan · 2021
Closest in time.
Instance-optimality in optimal value estimation: Adaptivity via variance-reduced Q-learning
Koulik Khamaru, Eric Xia, Martin J Wainwright, and Michael I Jordan · 2021
Closest in time.
Accelerated and instance-optimal policy evaluation with linear function approximation
Tianjiao Li, Guanghui Lan, and Ashwin Pananjady · 2021
Closest in time.
Instance-dependent ℓ ∞ \ell_{\infty} -bounds for policy evaluation in tabular reinforcement learning
Ashwin Pananjady and Martin J. Wainwright · 2021
Closest in time.
Statistical estimation of ergodic Markov chain kernel over discrete state space
Geoffrey Wolfer and Aryeh Kontorovich · 2021
Closest in time.