Fetching the paper…
Reading the bibliography…
We study the problem of estimating the fixed point of a contractive operator defined on a separable Banach space.
“A stochastic approximation method”
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
“Stochastic estimation of the maximum of a regression function”
Jack Kiefer and Jacob Wolfowitz · 1952
Earlier work this paper cites.
“On some asymptotic properties of maximum likelihood estimates and related Bayes estimates”
Lucien LeCam · 1953
Earlier work this paper cites.
“Denumerable state Markovian decision processes-average cost criterion”
Cyrus Derman · 1966
Earlier work this paper cites.
“On stochastic processes defined by differential equations with a small parameter”
Rafail Khas’minskii · 1966
Earlier work this paper cites.
“Local asymptotic minimax and admissibility in estimation”
Jaroslav H“”ajek · 1972
Earlier work this paper cites.
“Analysis of recursive stochastic algorithms”
Lennart Ljung · 1977
Earlier work this paper cites.
“On positive real transfer functions and the convergence of some recursive schemes”
Lennart Ljung · 1977
Earlier work this paper cites.
“Stochastic approximation methods for constrained and unconstrained systems” 26
Harold. Kushner and Dean. Clark · 1978
Earlier work this paper cites.
“Problem Complexity and Method Efficiency in Optimization”
Arkadic Nemirovski and David Yudin · 1983
Earlier work this paper cites.
“An invariant measure approach to the convergence of stochastic approximations with state dependent noise”
Harold Kushner and Adam Shwartz · 1984
Earlier work this paper cites.
“Approximation and weak convergence methods for random processes, with applications to stochastic systems theory”
Harold Kushner · 1984
Earlier work this paper cites.
“Efficient estimations from a slowly convergent Robbins-Monro process”, 1988
David Ruppert · 1988
Earlier work this paper cites.
“Recursive Methods in Economic Dynamics”
Nancy Stokey · 1989
Earlier work this paper cites.
“A new method of stochastic approximation type”
Boris Polyak · 1990
Earlier work this paper cites.
“Solving H-horizon, stationary Markov decision problems in time proportional to log ( H ) \log({H}) ”
Paul Tseng · 1990
Earlier work this paper cites.
“An analysis of stochastic shortest path problems”
Dimitri Bertsekas and John Tsitsiklis · 1991
Earlier work this paper cites.
“Acceleration of stochastic approximation by averaging”
Boris Polyak and Anatoli Juditsky · 1992
Earlier work this paper cites.
“Q-learning”
Christopher Watkins and Peter Dayan · 1992
Earlier work this paper cites.
“A dynamical system approach to stochastic approximations”
Michel Benaim · 1996
Earlier work this paper cites.
“New concentration inequalities in product spaces”
Michel Talagrand · 1996
Earlier work this paper cites.
“Stochastic and shortest path games: theory and algorithms”, 1997
Stephen Patek · 1997
Earlier work this paper cites.
“The asymptotic convergence-rate of Q-learning”
Csaba Szepesv“’ari · 1998
Earlier work this paper cites.
“Average cost temporal-difference learning”
John Tsitsiklis and Benjamin Van · 1999
Earlier work this paper cites.
“Asymptotic Statistics”
Aad van Vaart · 2000
Earlier work this paper cites.
“Finite-sample analysis of stochastic approximation using smooth convex envelopes”
Zaiwei Chen, Siva Theja, Sanjay Shakkottai and Karthikeyan Shanmugam · 2002
Earlier work this paper cites.
“Fixed Point Theory”
James Dugundji and Andrzej Granas · 2003
Earlier work this paper cites.
“Stochastic Approximation and Recursive Algorithms and Applications”
Harold Kushner and G Yin · 2003
Earlier work this paper cites.
“Introductory lectures on convex optimization: A basic course”
Yurii Nesterov · 2003
Cited alongside, same era.
“Markov decision processes: Discrete stochastic dynamic programming”
M.. Puterman · 2005
Cited alongside, same era.
“The Generic Chaining: Upper and Lower Bounds of Stochastic Processes”
Michel Talagrand · 2006
Cited alongside, same era.
“Stochastic Approximation: A Dynamical Systems Viewpoint”
Vivek Borkar · 2009
Cited alongside, same era.
“Robust stochastic approximation approach to stochastic programming”
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan and Alexander Shapiro · 2009
Cited alongside, same era.
“Variational analysis”
R Rockafellar and Roger J-B Wets · 2009
Cited alongside, same era.
“Probabilistic contraction analysis of iterated random operators”
Abhishek Gupta, Rahul Jain and Peter Glynn · 2018
Later among the works it cites.
Aaron Sidford et al · 2018
Later among the works it cites.
“Averaging stochastic gradient descent on Riemannian manifolds”
Nilesh Tripuraneni, Nicolas Flammarion, Francis Bach and Michael Jordan · 2018
Later among the works it cites.
“Reinforcement Learning and Optimal Control”
Dimitri Bertsekas · 2019
Later among the works it cites.
“Performance of Q-learning with linear function approximation: Stability and finite-time analysis”
Zaiwei Chen et al · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Solving variational inequalities with stochastic mirror-prox algorithm”
Anatoli Juditsky, Arkadi Nemirovski and Claire Tauvel · 2011
Cited alongside, same era.
“An Introduction to the Mathematical Theory of Inverse Problems”
Andreas Kirsch · 2011
Cited alongside, same era.
“Non-asymptotic analysis of stochastic approximation algorithms for machine learning”
“’Eric Moulines and Francis Bach · 2011
Cited alongside, same era.
“Approximate Dynamic Programming”
Dimitri Bertsekas · 2012
Cited alongside, same era.
“Weighted sup-norm contractions in dynamic programming: A review and some new applications”
Dimitri Bertsekas · 2012
Cited alongside, same era.
“Adaptive Algorithms and Stochastic Approximations”
Albert Benveniste, Michel M“’etivier and Pierre Priouret · 2012
Cited alongside, same era.
“Momentum-based variance reduction in non-convex SGD”
Ashok Cutkosky and Francesco Orabona · 2019
Later among the works it cites.
“High-dimensional Statistics: A Non-asymptotic Viewpoint”
Martin Wainwright · 2019
Later among the works it cites.
Martin Wainwright · 2019
Later among the works it cites.
“Variance-reduced Q-learning is minimax optimal”
Martin Wainwright · 2019
Later among the works it cites.
“Statistical inference for model parameters in stochastic gradient descent”
Xi Chen, Jason Lee, Xin Tong and Yichen Zhang · 2020
Later among the works it cites.
“Efficiently solving MDPs with stochastic mirror descent”
Yujia Jin and Aaron Sidford · 2020
Later among the works it cites.
“Is temporal difference learning optimal? An instance-dependent analysis”
Koulik Khamaru et al · 2020
Later among the works it cites.
Georgios Kotsalis, Guanghui Lan and Tianjiao Li · 2020
Later among the works it cites.
“Root-SGD: Sharp nonasymptotics and asymptotic efficiency in a single algorithm”
Chris Li, Wenlong Mou, Martin Wainwright and Michael Jordan · 2020
Later among the works it cites.
“On linear stochastic approximation: Fine-grained Polyak-Ruppert and non-asymptotic concentration”
Wenlong Mou et al · 2020
Later among the works it cites.
“Finite-Time Analysis of Asynchronous Stochastic Approximation and Q-Learning”
Guannan Qu and Adam Wierman · 2020
Later among the works it cites.
“A concentration bound for contractive stochastic approximation”
Vivek Borkar · 2021
Later among the works it cites.
“A Lyapunov theory for finite-sample guarantees of asynchronous Q-learning and TD-learning variants”
Zaiwei Chen, Siva Maguluri, Sanjay Shakkottai and Karthikeyan Shanmugam · 2021
Later among the works it cites.
“Finite-Sample Analysis of Off-Policy TD-Learning via Generalized Bellman Operators”
Zaiwei Chen, Siva Maguluri, Sanjay Shakkottai and Karthikeyan Shanmugam · 2021
Later among the works it cites.
“Instance-optimality in optimal value estimation: Adaptivity via variance-reduced Q-learning”
Koulik Khamaru, Eric Xia, Martin Wainwright and Michael Jordan · 2021
Later among the works it cites.
“Is Q-learning minimax optimal? a tight sample complexity analysis”
Gen Li et al · 2021
Later among the works it cites.
“Accelerated and instance-optimal policy evaluation with linear function approximation”
Tianjiao Li, Guanghui Lan and Ashwin Pananjady · 2021
Later among the works it cites.
“Optimal and instance-dependent guarantees for Markovian linear stochastic approximation”
Wenlong Mou, Ashwin Pananjady, Martin Wainwright and Peter Bartlett · 2021
Later among the works it cites.
“Inexact SARAH algorithm for stochastic optimization”
Lam Nguyen, Katya Scheinberg and Martin Tak“’ac · 2021
Later among the works it cites.
“Finite Sample Analysis of Average-Reward TD Learning and Q-Learning”
Sheng Zhang, Zhe Zhang and Siva Maguluri · 2021
Later among the works it cites.
“Instance-dependent confidence and early stopping for reinforcement learning”
Eric Xia, Koulik Khamaru, Martin. Wainwright and Michael. Jordan · 2022
Closest in time.
“Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization, II: shrinking procedures and optimal algorithms”
Saeed Ghadimi and Guanghui Lan · 2089
Closest in time.