Fetching the paper…
Reading the bibliography…
It is not clear yet why ADAM-alike adaptive gradient algorithms suffer from worse generalization performance than SGD despite their faster training speed.
Note on the derivatives with respect to a parameter of the solutions of a system of differential equations
T. Gronwall · 1919
Earlier work this paper cites.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Théorie de l’addition des variables aléatoires, gauthier-villars, paris, 1937
P. Levy · 1954
Earlier work this paper cites.
Optimal stabilization policies for deterministic and stochastic linear economic systems
Stephen J Turnovsky · 1973
Earlier work this paper cites.
Random integral equations with applications to life sciences and engineering
Chris P Tsokos and William J Padgett · 1974
Earlier work this paper cites.
Random differential equations as models of ecosystems: Monte carlo simulation approach
Jawahar Lal Tiwari and John E Hobbie · 1976
Earlier work this paper cites.
Random differential equations in river water quality modeling
Brad A Finney, David S Bowles, and Michael P Windham · 1982
Earlier work this paper cites.
Lectures on geometric measure theory
L. Simon · 1983
Earlier work this paper cites.
Stochastic gradient learning in neural networks
L. Bottou · 1991
Earlier work this paper cites.
Flat minima
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Stochastic differential equations with random coefficients
A. Kohatsu-Higa, J. León, and D. Nualart · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
S. Amari · 1998
Earlier work this paper cites.
Gradient based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Stabilization of continuous-time jump linear systems
Y. Fang and K. Loparo · 2002
Earlier work this paper cites.
Mean-variance portfolio selection with random parameters in a complete market
Andrew EB Lim and Xun Yu Zhou · 2002
Earlier work this paper cites.
Foundations of modern probability
O. Kallenberg · 2006
Earlier work this paper cites.
Cooling down Lévy flights
I. Pavlyukevich · 2007
Earlier work this paper cites.
An introduction to lévy processes with applications in finance
A. Papapantoleon · 2008
Earlier work this paper cites.
Learning deep architectures for AI
Y. Bengio · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky and G. Hinton · 2009
Earlier work this paper cites.
The hierarchy of exit times of lévy-driven langevin equations
P. Imkeller, I. Pavlyukevich, and T. Wetzel · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
First exit times of solutions of stochastic differential equations driven by multiplicative lévy noise with heavy tails
I. Pavlyukevich · 2011
Cited alongside, same era.
Deep neural networks for acoustic modeling in speech recognition
G. Hinton, L. Deng, D. Yu, G. Dahl, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, and B. Kingsbury · 2012
Cited alongside, same era.
Lecture 6.5—rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. Hinton · 2012
Cited alongside, same era.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Deep adversarial subspace clustering
P. Zhou, Y. Hou, and J. Feng · 2018
Later among the works it cites.
Averaging weights leads to wider optima and better generalization
P. Izmailov, D. Podoprikhin, T. Garipov, D. Vetrov, and A. Wilson · 2018
Later among the works it cites.
Visualizing the loss landscape of neural nets
H. Li, Z. Xu, G. Taylor, C. Studer, and T. Goldstein · 2018
Later among the works it cites.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
P. Chaudhari and S. Soatto · 2018
Later among the works it cites.
Efficient stochastic gradient hard thresholding
Pan Zhou, Xiaotong Yuan, and Jiashi Feng · 2018
Later among the works it cites.
New insight into hybrid stochastic gradient descent: Beyond with-replacement sampling and convexity
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Johnson and T. Zhang · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Cited alongside, same era.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Cited alongside, same era.
On estimating the tail index and the spectral measure of multivariate
M. Mohammadi, A. Mohammadpour, and H. Ogata · 2015
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
A variational analysis of stochastic gradient algorithms
S. Mandt, M. Hoffman, and D. Blei · 2016
Cited alongside, same era.
SGDR: Stochastic gradient descent with warm restarts
I. Loshchilov and F. Hutter · 2016
Cited alongside, same era.
P. Zhou, X. Yuan, and J. Feng · 2018
Later among the works it cites.
Understanding generalization and optimization performance of deep cnns
P. Zhou and J. Feng · 2018
Later among the works it cites.
Empirical risk landscape analysis for understanding deep neural networks
P. Zhou and J. Feng · 2018
Later among the works it cites.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
L. Wu, C. Ma, and W. E · 2018
Later among the works it cites.
Group normalization
Y. Wu and K. He · 2018
Later among the works it cites.
Asymmetric valleys: Beyond sharp and flat local minima
H. He, G. Huang, and Y. Yuan · 2019
Later among the works it cites.
Efficient meta learning via minibatch proximal update
P. Zhou, X. Yuan, H. Xu, S. Yan, and J. Feng · 2019
Later among the works it cites.
Theory-inspired path-regularized differential network architecture search
P. Zhou, C. Xiong, R. Socher, and S. Hoi · 2019
Later among the works it cites.
On the convergence of Adam and beyond
S. Reddi, S. Kale, and S. Kumar · 2019
Later among the works it cites.
Adaptive gradient methods with dynamic bound of learning rate
L. Luo, Y. Xiong, Y. Liu, and X. Sun · 2019
Later among the works it cites.
A tail-index analysis of stochastic gradient noise in deep neural networks
U. Simsekli, L. Sagun, and M. Gurbuzbalaban · 2019
Later among the works it cites.
Why adam beats sgd for attention models
J. Zhang, S. Karimireddy, A. Veit, S. Kim, S. Reddi, S. Kumar, and S. Sra · 2019
Later among the works it cites.
The anisotropic noise in stochastic gradient descent: Its behavior of escaping from minima and regularization effects
Z. Zhu, J. Wu, B. Yu, L. Wu, and J. Ma · 2019
Later among the works it cites.
Faster first-order methods for stochastic non-convex optimization on riemannian manifolds
P. Zhou, X. Yuan, and J. Feng · 2019
Later among the works it cites.
Gradient descent finds global minima of deep neural networks
S. Du, J. Lee, H. Li, L. Wang, and X. Zhai · 2019
Later among the works it cites.
Stability properties of systems of linear stochastic differential equations with random coefficients
A. Bishop and P. Del Moral · 2019
Later among the works it cites.
Hybrid stochastic-deterministic minibatch proximal gradient: Less-than-single-pass optimization with nearly optimal generalization
P. Zhou and X. Tong · 2020
Closest in time.