Fetching the paper…
Reading the bibliography…
Stochastic Gradient Descent with a constant learning rate (constant SGD) simulates a Markov chain with a stationary distribution.
Thermal agitation of electric charge in conductors
Harry Nyquist · 1928
Earlier work this paper cites.
On the theory of the Brownian motion
George E Uhlenbeck and Leonard S Ornstein · 1930
Earlier work this paper cites.
Brownian motion in a field of force and the diffusion model of chemical reactions
Hendrik Anthony Kramers · 1940
Earlier work this paper cites.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Boris Teodorovich Polyak · 1964
Earlier work this paper cites.
Efficient recursive estimation; application to estimating the parameters of a covariance function
David J Sakrison · 1965
Earlier work this paper cites.
Handbook of Stochastic Methods , volume 4
Crispin W Gardiner et al · 1985
Earlier work this paper cites.
Adaptive signal processing
Bernard Widrow and Samuel D Stearns · 1985
Earlier work this paper cites.
Asymptotic methods in statistical decision theory
Lucien Le Cam · 1986
Earlier work this paper cites.
A fast scoring algorithm for maximum likelihood estimation in unbalanced mixed models with nested random effects
Nicholas T Longford · 1987
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Boris T Polyak and Anatoli B Juditsky · 1992
Earlier work this paper cites.
Numerical methods for stochastic processes , volume 273
Nicolas Bouleau and Dominique Lepingle · 1994
Earlier work this paper cites.
Online learning and stochastic approximations
Léon Bottou · 1998
Earlier work this paper cites.
Introduction to variational methods for graphical models
Michael Jordan, Zoubin Ghahramani, Tommi Jaakkola, and Lawrence Saul · 1999
Earlier work this paper cites.
Propagation algorithms for variational Bayesian learning
Zoubin Ghahramani and Matthew J Beal · 2000
Earlier work this paper cites.
Advanced mean field methods: Theory and practice
Manfred Opper and David Saad · 2001
Earlier work this paper cites.
Stochastic approximation and recursive algorithms and applications , volume 35
Harold J Kushner and George Yin · 2003
Earlier work this paper cites.
Solving large scale linear prediction problems using stochastic gradient descent algorithms
Tong Zhang · 2004
Earlier work this paper cites.
Pattern Recognition and Machine Learning
Christopher Bishop · 2006
Cited alongside, same era.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Cited alongside, same era.
Bayesian learning via stochastic gradient Langevin dynamics
Max Welling and Yee W Teh · 2011
Cited alongside, same era.
Bayesian posterior sampling via stochastic gradient Fisher scoring
Sungjin Ahn, Anoop Korattikara, and Max Welling · 2012
Cited alongside, same era.
Stochastic approximation and optimization of random systems , volume 17
Lennart Ljung, Georg Ch Pflug, and Harro Walk · 2012
Cited alongside, same era.
Black box variational inference
Rajesh Ranganath, Sean Gerrish, and David M Blei · 2014
Later among the works it cites.
Stochastic backpropagation and approximate inference in deep generative models
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra · 2014
Later among the works it cites.
Approximation analysis of stochastic gradient Langevin dynamics by using Fokker-Planck equation and Ito process
Issei Sato and Hiroshi Nakagawa · 2014
Later among the works it cites.
Statistical analysis of stochastic gradient methods for generalized linear models
Panagiotis Toulis, Edoardo Airoldi, and Jason Rennie · 2014
Later among the works it cites.
On the convergence of stochastic gradient MCMC algorithms with high-order integrators
Changyou Chen, Nan Ding, and Lawrence Carin · 2015
Later among the works it cites.
Averaged least-mean-squares: Bias-variance trade-offs and optimal sampling distributions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lecture 6.5—RmsProp: Divide the Gradient by a Running Average of its Recent Magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Cited alongside, same era.
Non-strongly-convex smooth stochastic approximation with convergence rate o (1/n)
Francis Bach and Eric Moulines · 2013
Cited alongside, same era.
Stochastic variational inference
Matthew D Hoffman, David M Blei, Chong Wang, and John William Paisley · 2013
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Cited alongside, same era.
Fixed-form variational posterior approximation through stochastic linear regression
Tim Salimans and David A Knowles · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Cited alongside, same era.
Alexandre Défossez and Francis Bach · 2015
Later among the works it cites.
From averaging to acceleration, there is only a step-size
Nicolas Flammarion and Francis Bach · 2015
Later among the works it cites.
Automatic variational inference in STAN
Alp Kucukelbir, Rajesh Ranganath, Andrew Gelman, and David Blei · 2015
Later among the works it cites.
Dynamics of stochastic gradient algorithms
Qianxiao Li, Cheng Tai, and Weinan E · 2015
Later among the works it cites.
A complete recipe for stochastic gradient MCMC
Yi-An Ma, Tianqi Chen, and Emily B Fox · 2015
Later among the works it cites.
Covariance-controlled adaptive Langevin thermostat for large-scale Bayesian sampling
Xiaocheng Shang, Zhanxing Zhu, Benedict Leimkuhler, and Amos J Storkey · 2015
Later among the works it cites.
Bridging the gap between stochastic gradient MCMC and stochastic optimization
Changyou Chen, David Carlson, Zhe Gan, Chunyuan Li, and Lawrence Carin · 2016
Later among the works it cites.
Early stopping is nonparametric variational inference
Dougal Maclaurin, David Duvenaud, and Ryan P Adams · 2016
Later among the works it cites.
The generalized reparameterization gradient
Francisco Ruiz, Michaelis Titsias, and David Blei · 2016
Later among the works it cites.
Towards stability and optimality in stochastic gradient descent
Panos Toulis, Dustin Tran, and Edoardo M Airoldi · 2016
Later among the works it cites.
Bridging the gap between constant step size stochastic gradient descent and markov chains
Aymeric Dieuleveut, Alain Durmus, and Francis Bach · 2017
Closest in time.
Stochastic modified equations and adaptive stochastic gradient algorithms
Qianxiao Li, Cheng Tai, and Weinan E · 2017
Closest in time.