Fetching the paper…
Reading the bibliography…
Stochastic approximation (SA) is a key method used in statistical learning.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Maximum likelihood from incomplete data via the EM algorithm
Arthur P Dempster, Nan M Laird, and Donald B Rubin · 1977
Earlier work this paper cites.
On the convergence properties of the EM algorithm
CF Jeff Wu · 1983
Earlier work this paper cites.
Adaptive Algorithms and Stochastic Approximation
Albert Benveniste, Pierre Priouret, and Michel Métivier · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Convergence of stochastic iterative dynamic programming algorithms
Tommi Jaakkola, Michael I Jordan, and Satinder P Singh · 1994
Earlier work this paper cites.
Stochastic approximation with two time scales
Vivek S Borkar · 1997
Earlier work this paper cites.
Online learning and stochastic approximations
Léon Bottou · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Jonathan Baxter and Peter L Bartlett · 2001
Earlier work this paper cites.
On actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2003
Earlier work this paper cites.
Stochastic approximation and recursive algorithms and applications , volume 35
Harold Kushner and G George Yin · 2003
Earlier work this paper cites.
Natural actor-critic
Jan Peters and Stefan Schaal · 2008
Earlier work this paper cites.
Stochastic approximation: a dynamical systems viewpoint , volume 48
Vivek S Borkar · 2009
Cited alongside, same era.
On-line Expectation Maximization algorithm for latent data models
Olivier Cappé and Eric Moulines · 2009
Cited alongside, same era.
Convergence of adaptive and interacting Markov chain monte carlo algorithms
Gersende Fort, Eric Moulines, and Pierre Priouret · 2011
Cited alongside, same era.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Eric Moulines and Francis R Bach · 2011
Cited alongside, same era.
Thomas Degris, Martha White, and Richard S Sutton · 2012
Cited alongside, same era.
Ergodic mirror descent
John C Duchi, Alekh Agarwal, Mikael Johansson, and Michael I Jordan · 2012
Statistical guarantees for the EM algorithm: From population to sample-based analysis
Sivaraman Balakrishnan, Martin J Wainwright, Bin Yu, et al · 2017
Later among the works it cites.
Asymptotic bias of stochastic gradient search
Vladislav B Tadić and Arnaud Doucet · 2017
Later among the works it cites.
Regret bounds for model-free linear quadratic control
Yasin Abbasi-Yadkori, Nevena Lazic, and Csaba Szepesvari · 2018
Later among the works it cites.
A finite time analysis of temporal difference learning with linear function approximation
Jalaj Bhandari, Daniel Russo, and Raghav Singal · 2018
Later among the works it cites.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Later among the works it cites.
Stochastic Expectation Maximization with variance reduction
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The generalization ability of online algorithms for dependent data
Alekh Agarwal and John C Duchi · 2013
Cited alongside, same era.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Cited alongside, same era.
Nonlinear Time Series: Theory, Methods and Applications with R examples
Randal Douc, Eric Moulines, and David Stoffer · 2014
Cited alongside, same era.
Incremental majorization-minimization optimization with application to large-scale machine learning
Julien Mairal · 2015
Cited alongside, same era.
High dimensional em algorithm: Statistical optimization and asymptotic normality
Zhaoran Wang, Quanquan Gu, Yang Ning, and Han Liu · 2015
Cited alongside, same era.
Global analysis of Expectation Maximization for mixtures of two gaussians
Ji Xu, Daniel J Hsu, and Arian Maleki · 2016
Cited alongside, same era.
Jianfei Chen, Jun Zhu, Yee Whye Teh, and Tong Zhang · 2018
Later among the works it cites.
Spider: Near-optimal non-convex optimization via stochastic path-integrated differential estimator
Cong Fang, Chris Junchi Li, Zhouchen Lin, and Tong Zhang · 2018
Later among the works it cites.
Global convergence of policy gradient methods for the linear quadratic regulator
Maryam Fazel, Rong Ge, Sham Kakade, and Mehran Mesbahi · 2018
Later among the works it cites.
Linear stochastic approximation: How far does constant step-size and iterate averaging go?
Chandrashekar Lakshminarayanan and Csaba Szepesvari · 2018
Later among the works it cites.
Stochastic variance-reduced policy gradient
Matteo Papini, Damiano Binaghi, Giuseppe Canonaco, Matteo Pirotta, and Marcello Restelli · 2018
Later among the works it cites.
On Markov chain gradient descent
Tao Sun, Yuejiao Sun, and Wotao Yin · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction, 2nd Edition
Richard Sutton and Andrew Barto · 2018
Later among the works it cites.