Fetching the paper…
Reading the bibliography…
In this paper, we study the implicit regularization of stochastic gradient descent (SGD) through the lens of {\em dynamical stability} (Wu et al., 2018).
Universal approximation bounds for superpositions of a sigmoidal function
Barron, A. R · 1993
Earlier work this paper cites.
Simplifying neural nets by discovering flat minima
Hochreiter, S. and Schmidhuber, J · 1994
Earlier work this paper cites.
Flat minima
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Some PAC-Bayesian theorems
McAllester, D. A · 1999
Earlier work this paper cites.
Local Rademacher complexities
Bartlett, P. L., Bousquet, O., and Mendelson, S · 2005
Earlier work this paper cites.
Smoothness, low noise and fast rates
Srebro, N., Sridharan, K., and Tewari, A · 2010
Earlier work this paper cites.
Statistical language models based on neural networks
Mikolov, T. et al · 2012
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Pascanu, R., Mikolov, T., and Bengio, Y · 2013
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Neyshabur, B., Tomioka, R., and Srebro, N · 2014
Earlier work this paper cites.
Analysis of boolean functions
O’Donnell, R · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shalev-Shwartz, S. and Ben-David, S · 2014
Earlier work this paper cites.
Averaged least-mean-squares: Bias-variance trade-offs and optimal sampling distributions
Défossez, A. and Bach, F · 2015
Earlier work this paper cites.
Norm-based capacity control in neural networks
Neyshabur, B., Tomioka, R., and Srebro, N · 2015
Earlier work this paper cites.
Sharp minima can generalize for deep nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y · 2017
Earlier work this paper cites.
Three factors influencing minima in SGD
Jastrzębski, S., Kenton, Z., Arpit, D., Ballas, N., Fischer, A., Bengio, Y., and Storkey, A · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2017
Earlier work this paper cites.
Exploring generalization in deep learning
Neyshabur, B., Bhojanapalli, S., Mcallester, D., and Srebro, N · 2017
Earlier work this paper cites.
Towards understanding generalization of deep learning: Perspective of loss landscapes
Wu, L., Zhu, Z., and E, W · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2017
Cited alongside, same era.
Averaging weights leads to wider optima and better generalization
Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D., and Wilson, A. G · 2018
Cited alongside, same era.
High-dimensional probability: An introduction with applications in data science , volume 47
Vershynin, R · 2018
Cited alongside, same era.
How SGD selects the global minima in over-parameterized learning: A dynamical stability perspective
Wu, L., Ma, C., and E, W · 2018
Cited alongside, same era.
Kernel and rich regimes in overparametrized models
Woodworth, B., Gunasekar, S., Lee, J. D., Moroshko, E., Savarese, P., Golan, I., Soudry, D., and Srebro, N · 2020
Later among the works it cites.
Label noise SGD provably prefers flat global minimizers
Damian, A., Ma, T., and Lee, J · 2021
Later among the works it cites.
The Barron space and the flow-induced function spaces for neural network models
E, W., Ma, C., and Wu, L · 2021
Later among the works it cites.
The inverse variance–flatness relation in stochastic gradient descent is critical for finding flat minima
Feng, Y. and Tu, Y · 2021
Later among the works it cites.
What happens after SGD reaches zero loss?–a mathematical framework
Li, Z., Wang, T., and Arora, S · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A priori estimates of the population risk for two-layer neural networks
E, W., Ma, C., and Wu, L · 2019
Cited alongside, same era.
The implicit bias of depth: How incremental learning drives generalization
Gissin, D., Shalev-Shwartz, S., and Daniely, A · 2019
Cited alongside, same era.
Fisher-Rao metric, geometry, and complexity of neural networks
Liang, T., Poggio, T., Rakhlin, A., and Stokes, J · 2019
Cited alongside, same era.
The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects
Zhu, Z., Wu, J., Yu, B., Wu, L., and Ma, J · 2019
Cited alongside, same era.
Implicit gradient regularization
Barrett, D. and Dherin, B · 2020
Cited alongside, same era.
Implicit regularization for deep neural networks driven by an Ornstein-Uhlenbeck like process
Blanc, G., Gupta, N., Valiant, G., and Valiant, P · 2020
Cited alongside, same era.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Chizat, L. and Bach, F · 2020
Cited alongside, same era.
Noise and fluctuation of finite learning rate stochastic gradient descent
Liu, K., Ziyin, L., and Ueda, M · 2021
Later among the works it cites.
On linear stability of SGD and input-smoothness of neural networks
Ma, C. and Ying, L · 2021
Later among the works it cites.
The implicit bias of minima stability: A view from function space
Mulayoff, R., Michaeli, T., and Soudry, D · 2021
Later among the works it cites.
Implicit bias of SGD for diagonal linear networks: a provable benefit of stochasticity
Pesme, S., Pillaud-Vivien, L., and Flammarion, N · 2021
Later among the works it cites.
Relative flatness and generalization
Petzka, H., Kamp, M., Adilova, L., Sminchisescu, C., and Boley, M · 2021
Later among the works it cites.
Neurashed: A phenomenological model for imitating deep learning training
Su, W · 2021
Later among the works it cites.
Stochastic gradient descent with noise of machine learning type. part II: Continuous time analysis
Wojtowytsch, S · 2021
Later among the works it cites.
Towards understanding the condensation of two-layer neural networks at initial training
Xu, Z.-Q. J., Zhou, H., Luo, T., and Zhang, Y · 2021
Later among the works it cites.
Power-law escape rate of SGD
Mori, T., Ziyin, L., Liu, K., and Ueda, M · 2022
Later among the works it cites.
Implicit bias of the step size in linear diagonal neural networks
Nacson, M. S., Ravichandran, K., Srebro, N., and Soudry, D · 2022
Later among the works it cites.
On the implicit bias in deep-learning algorithms
Vardi, G · 2022
Later among the works it cites.
The alignment property of SGD noise and how it helps select flat minima: A stability analysis
Wu, L., Wang, M., and Su, W. J · 2022
Later among the works it cites.