Fetching the paper…
Reading the bibliography…
The noise in stochastic gradient descent (SGD), caused by minibatch sampling, is poorly understood despite its practical importance in deep learning.
Surprises in high-dimensional ridgeless least squares interpolation
Hastie, T., Montanari, A., Rosset, S., and Tibshirani, R. J. (2019) · 1903
Earlier work this paper cites.
Limitations of the empirical fisher approximation for natural gradient descent
Kunstner, F., Balles, L., and Hennig, P. (2019) · 1905
Earlier work this paper cites.
Numerical solution of the stable, non-negative definite lyapunov equation lyapunov equation
Hammarling, S. J. (1982) · 1982
Earlier work this paper cites.
On the expectation of the product of four matrix-valued gaussian random variables
Janssen, P. H. M. and Stoica, P. (1988) · 1988
Earlier work this paper cites.
The general problem of the stability of motion
Lyapunov, A. M. (1992) · 1992
Earlier work this paper cites.
Power laws are logarithmic boltzmann laws
Levy, M. and Solomon, S. (1996) · 1996
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, S.-I. (1998) · 1998
Earlier work this paper cites.
Stability theory for hybrid dynamical systems
Ye, H., Michel, A. N., and Hou, L. (1998) · 1998
Earlier work this paper cites.
The large learning rate phase of deep learning: the catapult mechanism
Lewkowycz, A., Bahri, Y., Dyer, E., Sohl-Dickstein, J., and Gur-Ari, G. (2020) · 2003
Earlier work this paper cites.
Shape matters: Understanding the implicit bias of the noise covariance
HaoChen, J. Z., Wei, C., Lee, J. D., and Ma, T. (2020) · 2006
Earlier work this paper cites.
Multiplicative noise and heavy tails in stochastic optimization
Hodgkinson, L. and Mahoney, M. W. (2020) · 2006
Earlier work this paper cites.
Dynamic of stochastic gradient descent with state-dependent noise
Meng, Q., Gong, S., Chen, W., Ma, Z.-M., and Liu, T.-Y. (2020) · 2006
Earlier work this paper cites.
Power-law distributions in empirical data
Clauset, A., Shalizi, C. R., and Newman, M. E. (2009) · 2009
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Welling, M. and Teh, Y. W. (2011) · 2011
Earlier work this paper cites.
Recent advances in deep learning theory
He, F. and Tao, D. (2020) · 2012
Earlier work this paper cites.
Noise and fluctuation of finite learning rate stochastic gradient descent
Liu, K., Ziyin, L., and Ueda, M. (2021) · 2012
Earlier work this paper cites.
New insights and perspectives on the natural gradient method
Martens, J. (2014) · 2014
Earlier work this paper cites.
Approximation analysis of stochastic gradient langevin dynamics by using fokker-planck equation and ito process
Sato, I. and Nakagawa, H. (2014) · 2014
Cited alongside, same era.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z. (2016) · 2016
Cited alongside, same era.
Computational methods for linear matrix equations
Simoncini, V. (2016) · 2016
Cited alongside, same era.
Second-order stochastic optimization for machine learning in linear time
Agarwal, N., Bullins, B., and Hazan, E. (2017) · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K. (2017) · 2017
Cited alongside, same era.
Fluctuation-dissipation relations for stochastic gradient descent
Yaida, S. (2019) · 2019
Later among the works it cites.
The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects
Zhu, Z., Wu, J., Yu, B., Wu, L., and Ma, J. (2019) · 2019
Later among the works it cites.
Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process
Blanc, G., Gupta, N., Valiant, G., and Valiant, P. (2020) · 2020
Later among the works it cites.
Bridging the gap between constant step size stochastic gradient descent and markov chains
Dieuleveut, A., Durmus, A., Bach, F., et al. (2020) · 2020
Later among the works it cites.
Uncertainty in neural networks: Approximately bayesian ensembling
Pearce, T., Leibfried, F., and Brintrup, A. (2020) · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Hoffer, E., Hubara, I., and Soudry, D. (2017) · 2017
Cited alongside, same era.
Stochastic gradient descent as approximate bayesian inference
Mandt, S., Hoffman, M. D., and Blei, D. M. (2017) · 2017
Cited alongside, same era.
Asymptotic and finite-sample properties of estimators based on stochastic gradients
Toulis, P., Airoldi, E. M., et al. (2017) · 2017
Cited alongside, same era.
A note on lazy training in supervised differentiable programming
Chizat, L. and Bach, F. (2018) · 2018
Cited alongside, same era.
Three factors influencing minima in SGD
Jastrzebski, S., Kenton, Z., Arpit, D., Ballas, N., Fischer, A., Storkey, A., and Bengio, Y. (2018) · 2018
Cited alongside, same era.
Xing, C., Arpit, D., Tsirigotis, C., and Bengio, Y. (2018) · 2018
Cited alongside, same era.
Theory of deep learning iib: Optimization properties of sgd
Zhang, C., Liao, Q., Rakhlin, A., Miranda, B., Golowich, N., and Poggio, T. (2018) · 2018
Cited alongside, same era.
Simsekli, U., Sener, O., Deligiannidis, G., and Erdogdu, M. A. (2020) · 2020
Later among the works it cites.
On the interplay between noise and curvature and its effect on optimization and generalization
Thomas, V., Pedregosa, F., Merriënboer, B., Manzagol, P.-A., Bengio, Y., and Le Roux, N. (2020) · 2020
Later among the works it cites.
Recent advances in deep learning
Wang, X., Zhao, Y., and Pourpanah, F. (2020) · 2020
Later among the works it cites.
On the noisy gradient descent that generalizes as sgd
Wu, J., Hu, W., Xiong, H., Huan, J., Braverman, V., and Zhu, Z. (2020) · 2020
Later among the works it cites.
Convergence rates and approximation results for sgd and its continuous-time counterpart
Fontaine, X., Bortoli, V. D., and Durmus, A. (2021) · 2021
Closest in time.
Kunin, D., Sagastuy-Brena, J., Gillespie, L., Margalit, E., Tanaka, H., Ganguli, S., and Yamins, D. L. (2021) · 2021
Closest in time.
Logarithmic landscape and power-law escape rate of sgd
Mori, T., Ziyin, L., Liu, K., and Ueda, M. (2021) · 2021
Closest in time.
Stochastic gradient descent with noise of machine learning type. part ii: Continuous time analysis
Wojtowytsch, S. (2021) · 2021
Closest in time.
A diffusion theory for deep learning dynamics: Stochastic gradient descent exponentially favors flat minima
Xie, Z., Sato, I., and Sugiyama, M. (2021) · 2021
Closest in time.
On the distributional properties of adaptive gradients
Zhang, Z. and Liu, Z. (2021) · 2021
Closest in time.
SGD can converge to local maxima
Ziyin, L., Li, B., Simon, J. B., and Ueda, M. (2022) · 2022
Closest in time.