Fetching the paper…
Reading the bibliography…
In the vanishing learning rate regime, stochastic gradient descent (SGD) is now relatively well understood.
Generalization in deep networks: The role of distance from initialization
Nagarajan, V. and Kolter, J. Z. (2019) · 1901
Earlier work this paper cites.
Über die von der molekularkinetischen theorie der wärme geforderte bewegung von in ruhenden flüssigkeiten suspendierten teilchen
Einstein, A. (1905) · 1905
Earlier work this paper cites.
Li, Y., Wei, C., and Ma, T. (2019) · 1907
Earlier work this paper cites.
Brownian motion in a field of force and the diffusion model of chemical reactions
Kramers, H. A. (1940) · 1940
Earlier work this paper cites.
Statistical physics, part 1
Landau, L. D. and Lifshitz, E. M. (1980) · 1980
Earlier work this paper cites.
A universal prior for integers and estimation by minimum description length
Rissanen, J. (1983) · 1983
Earlier work this paper cites.
Reaction-rate theory: fifty years after kramers
Hänggi, P., Talkner, P., and Borkovec, M. (1990) · 1990
Earlier work this paper cites.
A practical bayesian framework for backpropagation networks
MacKay, D. J. (1992) · 1992
Earlier work this paper cites.
Information and the accuracy attainable in the estimation of statistical parameters
Rao, C. R. (1992) · 1992
Earlier work this paper cites.
Stochastic processes in physics and chemistry
Van Kampen, N. G. (1992) · 1992
Earlier work this paper cites.
Flat minima
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, S.-I. (1998) · 1998
Earlier work this paper cites.
Online learning and stochastic approximations
Bottou, L. (1998) · 1998
Earlier work this paper cites.
Solving a quadratic matrix equation by newton’s method with exact line searches
Higham, N. J. and Kim, H.-M. (2001) · 2001
Earlier work this paper cites.
The large learning rate phase of deep learning: the catapult mechanism
Lewkowycz, A., Bahri, Y., Dyer, E., Sohl-Dickstein, J., and Gur-Ari, G. (2020) · 2003
Earlier work this paper cites.
Is deeper better? it depends on locality of relevant features
Mori, T. and Ueda, M. (2020b) · 2005
Earlier work this paper cites.
Multiplicative noise and heavy tails in stochastic optimization
Hodgkinson, L. and Mahoney, M. W. (2020) · 2006
Earlier work this paper cites.
Dynamic of stochastic gradient descent with state-dependent noise
Meng, Q., Gong, S., Chen, W., Ma, Z.-M., and Liu, T.-Y. (2020) · 2006
Earlier work this paper cites.
Methods of Information Geometry
Amari, S. and Nagaoka, H. (2007) · 2007
Earlier work this paper cites.
Improved generalization by noise enhancement
Mori, T. and Ueda, M. (2020a) · 2009
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y. (2011) · 2011
Cited alongside, same era.
Bayesian learning via stochastic gradient langevin dynamics
Welling, M. and Teh, Y. W. (2011) · 2011
Cited alongside, same era.
Efficient backprop
LeCun, Y. A., Bottou, L., Orr, G. B., and Müller, K.-R. (2012) · 2012
Cited alongside, same era.
Revisiting natural gradient for deep networks
Pascanu, R. and Bengio, Y. (2013) · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Quasi-hyperbolic momentum and adam for deep learning
Ma, J. and Yarats, D. (2018) · 2018
Later among the works it cites.
Lectures on convex optimization
Nesterov, Y. et al. (2018) · 2018
Later among the works it cites.
A bayesian perspective on generalization and stochastic gradient descent
Smith, S. L. and Le, Q. V. (2018) · 2018
Later among the works it cites.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
Wu, L., Ma, C., and E, W. (2018) · 2018
Later among the works it cites.
Theory of deep learning iib: Optimization properties of sgd
Zhang, C., Liao, Q., Rakhlin, A., Miranda, B., Golowich, N., and Poggio, T. (2018) · 2018
Later among the works it cites.
Entropy-sgd: Biasing gradient descent into wide valleys
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sutskever, I., Martens, J., Dahl, G., and Hinton, G. (2013) · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Cited alongside, same era.
From averaging to acceleration, there is only a step-size
Flammarion, N. and Bach, F. (2015) · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2015) · 2015
Cited alongside, same era.
A complete recipe for stochastic gradient mcmc
Ma, Y.-A., Chen, T., and Fox, E. (2015) · 2015
Cited alongside, same era.
Dziugaite, G. K. and Roy, D. M. (2017) · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K. (2017) · 2017
Cited alongside, same era.
Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J., Sagun, L., and Zecchina, R. (2019) · 2019
Later among the works it cites.
Understanding the role of momentum in stochastic gradient methods
Gitman, I., Lang, H., Zhang, P., and Xiao, L. (2019) · 2019
Later among the works it cites.
On the diffusion approximation of nonconvex stochastic gradient descent
Hu, W., Li, C. J., Li, L., and Liu, J.-G. (2019) · 2019
Later among the works it cites.
Fantastic generalization measures and where to find them
Jiang, Y., Neyshabur, B., Mobahi, H., Krishnan, D., and Bengio, S. (2019) · 2019
Later among the works it cites.
Fisher-rao metric, geometry, and complexity of neural networks
Liang, T., Poggio, T., Rakhlin, A., and Stokes, J. (2019) · 2019
Later among the works it cites.
A tail-index analysis of stochastic gradient noise in deep neural networks
Simsekli, U., Sagun, L., and Gurbuzbalaban, M. (2019) · 2019
Later among the works it cites.
The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects
Zhu, Z., Wu, J., Yu, B., Wu, L., and Ma, J. (2019) · 2019
Later among the works it cites.
Bridging the gap between constant step size stochastic gradient descent and markov chains
Dieuleveut, A., Durmus, A., Bach, F., et al. (2020) · 2020
Closest in time.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J., Xiao, L., Schoenholz, S. S., Bahri, Y., Novak, R., Sohl-Dickstein, J., and Pennington, J. (2020) · 2020
Closest in time.
Hausdorff dimension, heavy tails, and generalization in neural networks
Simsekli, U., Sener, O., Deligiannidis, G., and Erdogdu, M. A. (2020) · 2020
Closest in time.
Logarithmic landscape and power-law escape rate of sgd
Mori, T., Ziyin, L., Liu, K., and Ueda, M. (2021) · 2021
Closest in time.
A diffusion theory for deep learning dynamics: Stochastic gradient descent exponentially favors flat minima
Xie, Z., Sato, I., and Sugiyama, M. (2021) · 2021
Closest in time.
On the distributional properties of adaptive gradients
Zhiyi, Z. and Ziyin, L. (2021) · 2021
Closest in time.
On minibatch noise: Discrete-time sgd, overparametrization, and bayes
Ziyin, L., Liu, K., Mori, T., and Ueda, M. (2021) · 2021
Closest in time.