Fetching the paper…
Reading the bibliography…
Diffusion approximation provides weak approximation for stochastic gradient descent algorithms in a finite time horizon.
A Stochastic Approximation Method
Herbert Robbins and Sutton Monro · 1985
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Boris T Polyak and Anatoli B Juditsky · 1992
Earlier work this paper cites.
Iterated random functions
Persi Diaconis and David Freedman · 1999
Earlier work this paper cites.
Stochastic differential equations: an introduction with applications
B. Øksendal · 2003
Earlier work this paper cites.
Topics in optimal transportation
Cédric Villani · 2003
Earlier work this paper cites.
Solving large scale linear prediction problems using stochastic gradient descent algorithms
Tong Zhang · 2004
Earlier work this paper cites.
Quasi-geodesic neural learning algorithms over the orthogonal group: A tutorial
Simone Fiori · 2005
Earlier work this paper cites.
Modified equations for stochastic differential equations
Tony Shardlow · 2006
Earlier work this paper cites.
Stochastic Convex Optimization
Shai Shalev-Shwartz, Ohad Shamir, Nathan Srebro, and Karthik Sridharan · 2009
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Eric Moulines and Francis R. Bach · 2011
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Alexander Rakhlin, Ohad Shamir, and Karthik Sridharan · 2012
Earlier work this paper cites.
Weak backward error analysis for SDEs
Arnaud Debussche and Erwan Faou · 2012
Earlier work this paper cites.
High weak order methods for stochastic differential equations based on modified equations
Assyr Abdulle, David Cohen, Gilles Vilmart, and Konstantinos C Zygalakis · 2012
Earlier work this paper cites.
Optimization and dynamical systems
Uwe Helmke and John B Moore · 2012
Earlier work this paper cites.
A smooth vector field for quadratic programming
Hans-Bernd Dörr, Erkin Saka, and Christian Ebenbauer · 2012
Cited alongside, same era.
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
Ohad Shamir and Tong Zhang · 2013
Cited alongside, same era.
Non-strongly-convex smooth stochastic approximation with convergence rate O ( 1 / n ) O(1/n)
Francis Bach and Eric Moulines · 2013
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Cited alongside, same era.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Cited alongside, same era.
High order numerical approximation of the invariant measure of ergodic SDEs
Improving generalization performance by switching from Adam to SGD
Nitish Shirish Keskar and Richard Socher · 2017
Later among the works it cites.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Later among the works it cites.
Bridging the gap between constant step size stochastic gradient descent and Markov chains
A. Dieuleveut, A. Durmus, and F. Bach · 2017
Later among the works it cites.
Stochastic modified equations and adaptive stochastic gradient algorithms
Qianxiao Li, Cheng Tai, and Weinan E · 2017
Later among the works it cites.
Stochastic gradient descent in continuous time
Justin Sirignano and Konstantinos Spiliopoulos · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Assyr Abdulle, Gilles Vilmart, and Konstantinos C Zygalakis · 2014
Cited alongside, same era.
Weak backward error analysis for overdamped langevin processes
Marie Kopec · 2014
Cited alongside, same era.
On variance reduction in stochastic gradient descent and its asynchronous variants
Sashank J Reddi, Ahmed Hefny, Suvrit Sra, Barnabas Poczos, and Alexander J Smola · 2015
Cited alongside, same era.
A globally convergent incremental newton method
Mert Gürbüzbalaban, Asuman Ozdaglar, and Pablo Parrilo · 2015
Cited alongside, same era.
Weak backward error analysis for langevin process
Marie Kopec · 2015
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, and Klaus Macherey · 2016
Cited alongside, same era.
Sparse recovery via differential inclusions
Stanley Osher, Feng Ruan, Jiechao Xiong, Yuan Yao, and Wotao Yin · 2016
Cited alongside, same era.
Stochastic gradient descent in continuous time: A central limit theorem
Justin Sirignano and Konstantinos Spiliopoulos · 2017
Later among the works it cites.
Prateek Jain, Sham M Kakade, Rahul Kidambi, Praneeth Netrapalli, Venkata Krishna Pillutla, and Aaron Sidford · 2017
Later among the works it cites.
Stochastic gradient descent as approximate bayesian inference
Stephan Mandt, Matthew D. Hoffman, and David M. Blei · 2017
Later among the works it cites.
Strong error analysis for stochastic gradient descent optimization algorithms
Arnulf Jentzen, Benno Kuckuck, Ariel Neufeld, and Philippe von Wurstemberger · 2018
Later among the works it cites.
Qianxiao Li, Cheng Tai, and E Weinan · 2018
Later among the works it cites.
Semi-groups of stochastic gradient descent and online principal component analysis: properties and diffusion approximations
Yuanyuan Feng, Lei Li, and Jian-Guo Liu · 2018
Later among the works it cites.
On the diffusion approximation of nonconvex stochastic gradient descent
W. Hu, C. J. Li, L. Li, and J.-G. Liu · 2018
Later among the works it cites.
Deep relaxation: partial differential equations for optimizing deep neural networks
Pratik Chaudhari, Adam Oberman, Stanley Osher, Stefano Soatto, and Guillaume Carlier · 2018
Later among the works it cites.
Adding one neuron can eliminate all bad local minima
Shiyu Liang, Ruoyu Sun, Jason D Lee, and R. Srikant · 2018
Later among the works it cites.