Fetching the paper…
Reading the bibliography…
Stochastic gradient descent (SGD) has been widely used in machine learning due to its computational efficiency and favorable generalization properties.
A method for simulating stable random variables
J. M. Chambers, C. L. Mallows, and B. W. Stuck · 1976
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Lectures on the coupling method
Torgny Lindvall · 2002
Earlier work this paper cites.
Financial modelling with jump processes
Peter Tankov · 2003
Earlier work this paper cites.
Metastability in reversible diffusion processes I: Sharp asymptotics for capacities and exit times
Anton Bovier, Michael Eckhoff, Véronique Gayrard, and Markus Klein · 2004
Earlier work this paper cites.
Applied stochastic control of jump diffusions
Bernt Karsten Øksendal and Agnes Sulem · 2005
Earlier work this paper cites.
First exit times of sdes driven by stable Lévy processes
Peter Imkeller and Ilya Pavlyukevich · 2006
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Earlier work this paper cites.
The hierarchy of exit times of Lévy-driven Langevin equations
P. Imkeller, I. Pavlyukevich, and T. Wetzel · 2010
Earlier work this paper cites.
First exit times of non-linear dynamical systems in rd perturbed by multifractal Lévy noise
Peter Imkeller, Ilya Pavlyukevich, and Michael Stauch · 2010
Earlier work this paper cites.
First exit times of solutions of stochastic differential equations driven by multiplicative Lévy noise with heavy tails
Ilya Pavlyukevich · 2011
Earlier work this paper cites.
Kramers’ law: Validity, derivations and generalisations
Nils Berglund · 2011
Earlier work this paper cites.
Stochastic gradient descent tricks
Léon Bottou · 2012
Earlier work this paper cites.
Pathwise uniqueness for singular sdes driven by stable processes
Enrico Priola et al · 2012
Earlier work this paper cites.
Moments and absolute moments of the normal distribution
Andreas Winkelbauer · 2012
Earlier work this paper cites.
Existence of densities for stable-like driven sde’s with Hölder continuous coefficients
Arnaud Debussche and Nicolas Fournier · 2013
Cited alongside, same era.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Cited alongside, same era.
An Introduction to Stochastic Dynamics
J. Duan · 2015
Cited alongside, same era.
Spectral Analysis for a Discrete Metastable System Driven by Lévy Flights
Toralf Burghoff and Ilya Pavlyukevich · 2015
Cited alongside, same era.
Weak reflection principle for Lévy processes
Erhan Bayraktar, Sergey Nadtochiy, et al · 2015
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2016
Ergodicity of stochastic differential equations with jumps and singular coefficients
Longjie Xie and Xicheng Zhang · 2017
Later among the works it cites.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
P. Chaudhari and S. Soatto · 2018
Later among the works it cites.
Zhanxing Zhu, Jingfeng Wu, Bing Yu, Lei Wu, and Jinwen Ma · 2018
Later among the works it cites.
Local Optimality and Generalization Guarantees for the Langevin Algorithm via Empirical Metastability
B. Tzen, T. Liang, and M. Raginsky · 2018
Later among the works it cites.
Gradient Estimates and Ergodicity for SDEs Driven by Multiplicative Lévy Noises via Coupling
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
A variational analysis of stochastic gradient algorithms
S. Mandt, M. Hoffman, and D. Blei · 2016
Cited alongside, same era.
Eigenvalues of the hessian in deep learning: Singularity and beyond
Levent Sagun, Leon Bottou, and Yann LeCun · 2016
Cited alongside, same era.
Three factors influencing minima in SGD
S. Jastrzebski, Z. Kenton, D. Arpit, N. Ballas, A. Fischer, Y. Bengio, and A. Storkey · 2017
Cited alongside, same era.
Stochastic Modified Equations and Adaptive Stochastic Gradient Algorithms
Q. Li, C. Tai, and W. E · 2017
Cited alongside, same era.
On the diffusion approximation of nonconvex stochastic gradient descent
W. Hu, C. J. Li, L. Li, and J.-G. Liu · 2017
Cited alongside, same era.
Mingjie Liang and Jian Wang · 2018
Later among the works it cites.
Global convergence of Langevin dynamics based algorithms for nonconvex optimization
Pan Xu, Jinghui Chen, Difan Zou, and Quanquan Gu · 2018
Later among the works it cites.
Global Non-convex Optimization with Discretized Diffusions
M. A. Erdogdu, L. Mackey, and O. Shamir · 2018
Later among the works it cites.
Breaking Reversibility Accelerates Langevin Dynamics for Global Non-Convex Optimization
Xuefeng Gao, Mert Gurbuzbalaban, and Lingjiong Zhu · 2018
Later among the works it cites.
Xuefeng Gao, Mert Gürbüzbalaban, and Lingjiong Zhu · 2018
Later among the works it cites.
The full spectrum of deep net hessians at scale: Dynamics with sample size
Vardan Papyan · 2018
Later among the works it cites.
On the rate of convergence of strong Euler approximation for SDEs driven by lévy processes
R Mikulevičius and Fanhui Xu · 2018
Later among the works it cites.
A Tail-Index Analysis of Stochastic Gradient Noise in Deep Neural Networks
U. Şimşekli, L. Sagun, and Gürbüzbalaban · 2019
Closest in time.
Non-Asymptotic Analysis of Fractional Langevin Monte Carlo for Non-Convex Optimization
Thanh Huy Nguyen, Umut Şimşekli, and Gaël Richard · 2019
Closest in time.
Fluctuation-dissipation relations for stochastic gradient descent
S. Yaida · 2019
Closest in time.
On weak uniqueness and distributional properties of a solution to an sde with α \alpha -stable noise
Alexei M Kulik · 2019
Closest in time.