Fetching the paper…
Reading the bibliography…
Adaptive Moment Estimation (Adam), which combines Adaptive Learning Rate and Momentum, would be the most popular stochastic optimizer for accelerating the training of deep neural networks.
The activated complex in chemical reactions
Eyring, H · 1935
Earlier work this paper cites.
Brownian motion in a field of force and the diffusion model of chemical reactions
Kramers, H. A · 1940
Earlier work this paper cites.
The fokker-planck equation, methods of solution and applications
Risken, H. and Eberly, J · 1985
Earlier work this paper cites.
Escape from a metastable state
Hanggi, P · 1986
Earlier work this paper cites.
Reaction-rate theory: fifty years after kramers
Hänggi, P., Talkner, P., and Borkovec, M · 1990
Earlier work this paper cites.
Stochastic processes in physics and chemistry , volume 1
Van Kampen, N. G · 1992
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Marcus, M., Santorini, B., and Marcinkiewicz, M. A · 1993
Earlier work this paper cites.
Heavy-ball method in nonconvex optimization problems
Zavriev, S. and Kostyuk, F · 1993
Earlier work this paper cites.
Simplifying neural nets by discovering flat minima
Hochreiter, S. and Schmidhuber, J · 1995
Earlier work this paper cites.
The effects of adding noise during backpropagation training on a generalization performance
An, G · 1996
Earlier work this paper cites.
Fokker-planck equation
Risken, H · 1996
Earlier work this paper cites.
In all likelihood: statistical modelling and inference using likelihood
Pawitan, Y · 2001
Earlier work this paper cites.
Elements of nonequilibrium statistical mechanics , volume 3
Balakrishnan, V · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Rate theories for biologists
Zhou, H.-X · 2010
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Welling, M. and Teh, Y. W · 2011
Earlier work this paper cites.
Stable weight decay regularization
Xie, Z., Sato, I., and Sugiyama, M · 2011
Earlier work this paper cites.
The Langevin equation: with applications to stochastic problems in physics, chemistry and electrical engineering , volume 27
Coffey, W. and Kalmykov, Y. P · 2012
Earlier work this paper cites.
Neural networks for machine learning lecture 6a overview of mini-batch gradient descent
Hinton, G., Srivastava, N., and Swersky, K · 2012
Earlier work this paper cites.
Phase space reduction of the one-dimensional fokker-planck (kramers) equation
Kalinay, P. and Percus, J. K · 2012
Earlier work this paper cites.
Kramers’ law: Validity, derivations and generalisations
Berglund, N · 2013
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Ghadimi, S. and Lan, G · 2013
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y. N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y · 2014
Earlier work this paper cites.
Approximation analysis of stochastic gradient langevin dynamics by using fokker-planck equation and ito process
Sato, I. and Nakagawa, H · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Cited alongside, same era.
Recurrent neural network regularization
Zaremba, W., Sutskever, I., and Vinyals, O · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Cited alongside, same era.
Adding gradient noise improves learning for very deep networks
Neelakantan, A., Vilnis, L., Le, Q. V., Sutskever, I., Kaiser, L., Kurach, K., and Martens, J · 2015
Cited alongside, same era.
Adaptive methods for nonconvex optimization
Zaheer, M., Reddi, S., Sachan, D., Kale, S., and Kumar, S · 2018
Later among the works it cites.
On the diffusion approximation of nonconvex stochastic gradient descent
Hu, W., Li, C. J., Li, L., and Liu, J.-G · 2019
Later among the works it cites.
Fantastic generalization measures and where to find them
Jiang, Y., Neyshabur, B., Mobahi, H., Krishnan, D., and Bengio, S · 2019
Later among the works it cites.
Stochastic modified equations and dynamics of stochastic gradient algorithms i: Mathematical foundations
Li, Q., Tai, C., and Weinan, E · 2019
Later among the works it cites.
On the variance of the adaptive learning rate and beyond
Liu, L., Jiang, H., He, P., Chen, W., Liu, X., Gao, J., and Han, J · 2019
Later among the works it cites.
Adaptive gradient methods with dynamic bound of learning rate
Luo, L., Xiong, Y., Liu, Y., and Sun, X · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2015
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Hardt, M., Recht, B., and Singer, Y · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Densely connected convolutional networks
Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q · 2017
Cited alongside, same era.
Three factors influencing minima in sgd
Jastrzkebski, S., Kenton, Z., Arpit, D., Ballas, N., Fischer, A., Bengio, Y., and Storkey, A · 2017
Cited alongside, same era.
How to escape saddle points efficiently
Jin, C., Ge, R., Netrapalli, P., Kakade, S. M., and Jordan, M. I · 2017
Cited alongside, same era.
Stochastic modified equations and adaptive stochastic gradient algorithms
Li, Q., Tai, C., et al · 2017
Cited alongside, same era.
Later among the works it cites.
First exit time analysis of stochastic gradient descent under heavy-tailed gradient noise
Nguyen, T. H., Simsekli, U., Gurbuzbalaban, M., and Richard, G · 2019
Later among the works it cites.
On the convergence of adam and beyond
Reddi, S. J., Kale, S., and Kumar, S · 2019
Later among the works it cites.
A tail-index analysis of stochastic gradient noise in deep neural networks
Simsekli, U., Sagun, L., and Gurbuzbalaban, M · 2019
Later among the works it cites.
Escaping saddle points with adaptive gradient methods
Staib, M., Reddi, S., Kale, S., Kumar, S., and Sra, S · 2019
Later among the works it cites.
Escaping saddle points faster with stochastic momentum
Wang, J.-K., Lin, C.-H., and Abernethy, J · 2019
Later among the works it cites.
Toward understanding the importance of noise in training neural networks
Zhou, M., Liu, T., Li, Y., Lin, D., Zhou, E., and Zhao, T · 2019
Later among the works it cites.
The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects
Zhu, Z., Wu, J., Yu, B., Wu, L., and Ma, J · 2019
Later among the works it cites.
A sufficient condition for convergences of adam and rmsprop
Zou, F., Shen, L., Jie, Z., Zhang, W., and Liu, W · 2019
Later among the works it cites.
On the convergence of adam and adagrad
Défossez, A., Bottou, L., Bach, F., and Usunier, N · 2020
Closest in time.
Langevin equation and fokker-planck equation
Radpay, P · 2020
Closest in time.
Stochastic gradient descent with nonlinear conjugate gradient-style adaptive momentum
Wang, B. and Ye, Q · 2020
Closest in time.
Towards theoretically understanding why sgd generalizes better than adam in deep learning
Zhou, P., Feng, J., Ma, C., Xiong, C., Hoi, S. C. H., et al · 2020
Closest in time.
Adabelief optimizer: Adapting stepsizes by the belief in observed gradients
Zhuang, J., Tang, T., Ding, Y., Tatikonda, S. C., Dvornek, N., Papademetris, X., and Duncan, J · 2020
Closest in time.
Shape matters: Understanding the implicit bias of the noise covariance
HaoChen, J. Z., Wei, C., Lee, J., and Ma, T · 2021
Closest in time.
On the validity of modeling sgd with stochastic differential equations (sdes)
Li, Z., Malladi, S., and Arora, S · 2021
Closest in time.
{RMS}prop can converge with proper hyper-parameter
Shi, N., Li, D., Hong, M., and Sun, R · 2021
Closest in time.
On the power-law spectrum in deep learning: A bridge to protein science
Xie, Z., Tang, Q.-Y., Cai, Y., Sun, M., and Li, P · 2022
Closest in time.