Fetching the paper…
Reading the bibliography…
We showcase important features of the dynamics of the Stochastic Gradient Descent (SGD) in the training of neural networks.
Stochastic differential equations: an introduction with applications
Oksendal, B · 2013
Earlier work this paper cites.
Exploiting linear structure within convolutional networks for efficient evaluation
Denton, E. L., Zaremba, W., Bruna, J., LeCun, Y., and Fergus, R · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., Dean, J., et al · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Earlier work this paper cites.
Sharp minima can generalize for deep nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2017
Earlier work this paper cites.
Understanding batch normalization
Bjorck, N., Gomes, C. P., Selman, B., and Weinberger, K. Q · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A., Gabriel, F., and Hongler, C · 2018
Earlier work this paper cites.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Li, Y., Ma, T., and Zhang, H · 2018
Earlier work this paper cites.
A Bayesian perspective on generalization and stochastic gradient descent
Smith, S. L. and Le, Q. V · 2018
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Soudry, D., Hoffer, E., Nacson, M. S., Gunasekar, S., and Srebro, N · 2018
Earlier work this paper cites.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
Wu, L., Ma, C., et al · 2018
Earlier work this paper cites.
Xing, C., Arpit, D., Tsirigotis, C., and Bengio, Y · 2018
Earlier work this paper cites.
Three mechanisms of weight decay regularization
Zhang, G., Wang, C., Xu, B., and Grosse, R · 2018
Earlier work this paper cites.
On lazy training in differentiable programming
Chizat, L., Oyallon, E., and Bach, F · 2019
Earlier work this paper cites.
An exponential learning rate schedule for deep learning
Li, Z. and Arora, S · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Cited alongside, same era.
Implicit regularization for optimal sparse recovery
Vaskevicius, T., Kanade, V., and Rebeschini, P · 2019
Cited alongside, same era.
Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process
Blanc, G., Gupta, N., Valiant, G., and Valiant, P · 2020
Cited alongside, same era.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Chizat, L. and Bach, F · 2020
Cited alongside, same era.
The large learning rate phase of deep learning: the catapult mechanism
Catastrophic fisher explosion: Early phase fisher matrix impacts generalization
Jastrzebski, S., Arpit, D., Astrand, O., Kerg, G. B., Wang, H., Xiong, C., Socher, R., Cho, K., and Geras, K. J · 2021
Later among the works it cites.
The Sobolev regularization effect of stochastic gradient descent
Ma, C. and Ying, L · 2021
Later among the works it cites.
The implicit bias of minima stability: A view from function space
Mulayoff, R., Michaeli, T., and Soudry, D · 2021
Later among the works it cites.
Implicit bias of sgd for diagonal linear networks: a provable benefit of stochasticity
Pesme, S., Pillaud-Vivien, L., and Flammarion, N · 2021
Later among the works it cites.
On the origin of implicit regularization in stochastic gradient descent
Smith, S. L., Dherin, B., Barrett, D. G., and De, S · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lewkowycz, A., Bahri, Y., Dyer, E., Sohl-Dickstein, J., and Gur-Ari, G · 2020
Cited alongside, same era.
Reconciling modern deep learning with traditional optimization analyses: The intrinsic learning rate
Li, Z., Lyu, K., and Arora, S · 2020
Cited alongside, same era.
Gradient descent maximizes the margin of homogeneous neural networks
Lyu, K. and Li, J · 2020
Cited alongside, same era.
Implicit bias in deep linear classification: Initialization scale vs training accuracy
Moroshko, E., Woodworth, B. E., Gunasekar, S., Lee, J. D., Srebro, N., and Soudry, D · 2020
Cited alongside, same era.
Learning rate annealing can provably help generalization, even for convex problems
Nakkiran, P · 2020
Cited alongside, same era.
Kernel and rich regimes in overparametrized models
Woodworth, B., Gunasekar, S., Lee, J. D., Moroshko, E., Savarese, P., Golan, I., Soudry, D., and Srebro, N · 2020
Cited alongside, same era.
Acceleration via fractal learning rate schedules
Agarwal, N., Goel, S., and Zhang, C · 2021
Cited alongside, same era.
Direction matters: On the implicit bias of stochastic gradient descent with moderate learning rate
Wu, J., Zou, D., Braverman, V., and Gu, Q · 2021
Later among the works it cites.
On the benefits of large learning rates for kernel methods
Beugnot, G., Mairal, J., and Rudi, A · 2022
Closest in time.
Gradient flow dynamics of shallow relu networks for square loss and orthogonal inputs
Boursier, E., Pillaud-Vivien, L., and Flammarion, N · 2022
Closest in time.
On gradient descent convergence beyond the edge of stability
Chen, L. and Bruna, J · 2022
Closest in time.
Stochastic training is not necessary for generalization
Geiping, J., Goldblum, M., Pope, P. E., Moeller, M., and Goldstein, T · 2022
Closest in time.
On the maximum Hessian eigenvalue and generalization
Kaur, S., Cohen, J., and Lipton, Z. C · 2022
Closest in time.
What happens after sgd reaches zero loss?–a mathematical framework
Li, Z., Wang, T., and Arora, S · 2022
Closest in time.
The multiscale structure of neural network loss functions: The effect on optimization and origin
Ma, C., Wu, L., and Ying, L · 2022
Closest in time.
Implicit bias of the step size in linear diagonal neural networks
Nacson, M. S., Ravichandran, K., Srebro, N., and Soudry, D · 2022
Closest in time.
Label noise (stochastic) gradient descent implicitly solves the lasso for quadratic parametrisation
Pillaud-Vivien, L., Reygner, J., and Flammarion, N · 2022
Closest in time.
Large learning rate tames homogeneity: Convergence and balancing effect
Wang, Y., Chen, M., Zhao, T., and Tao, M · 2022
Closest in time.
Yang, N., Tang, C., and Tu, Y · 2022
Closest in time.
Strength of minibatch noise in sgd
Ziyin, L., Liu, K., Mori, T., and Ueda, M · 2022
Closest in time.