Fetching the paper…
Reading the bibliography…
The largest eigenvalue of the Hessian, or sharpness, of neural networks is a key quantity to understand their optimization dynamics.
Sur une manière d’étendre le théorème de la moyenne aux équations différentielles du premier ordre
M. Michel Petrovitch · 1901
Earlier work this paper cites.
Flat minima
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
A. Saxe, J. McClelland, and S. Ganguli · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Gaussian error linear units (gelus)
D. Hendrycks and K. Gimpel · 2016
Earlier work this paper cites.
Sharp minima can generalize for deep nets
L. Dinh, R. Pascanu, S. Bengio, and Y. Bengio · 2017
Earlier work this paper cites.
Implicit regularization in matrix factorization
S. Gunasekar, B. E. Woodworth, S. Bhojanapalli, B. Neyshabur, and N. Srebro · 2017
Earlier work this paper cites.
Three factors influencing minima in SGD
S. Jastrzębski, Z. Kenton, D. Arpit, N. Ballas, A. Fischer, Y. Bengio, and A. Storkey · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2017
Earlier work this paper cites.
Exploring generalization in deep learning
B. Neyshabur, S. Bhojanapalli, D. Mcallester, and N. Srebro · 2017
Earlier work this paper cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
S. Arora, N. Cohen, and E. Hazan · 2018
Earlier work this paper cites.
Gradient descent with identity initialization efficiently learns positive definite linear transformations by deep residual networks
P. Bartlett, D. Helmbold, and P. Long · 2018
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. VanderPlas, S. Wanderman-Milne, and Q. Zhang · 2018
Earlier work this paper cites.
Lectures on Convex Optimization
Y. Nesterov · 2018
Earlier work this paper cites.
A bayesian perspective on generalization and stochastic gradient descent
S. L. Smith and Q. V. Le · 2018
Earlier work this paper cites.
High-Dimensional Probability: An Introduction with Applications in Data Science
R. Vershynin · 2018
Earlier work this paper cites.
How to initialize your network? Robust initialization for WeightNorm & ResNets
D. Arpit, V. Campos, and Y. Bengio · 2019
Cited alongside, same era.
On lazy training in differentiable programming
L. Chizat, E. Oyallon, and F. Bach · 2019
Cited alongside, same era.
Implicit regularization of discrete gradient dynamics in linear neural networks
G. Gidel, F. Bach, and S. Lacoste-Julien · 2019
Cited alongside, same era.
An analytic theory of generalization dynamics and transfer learning in deep linear networks
A. K. Lampinen and S. Ganguli · 2019
Cited alongside, same era.
A mathematical theory of semantic development in deep neural networks
A. M. Saxe, J. L. McClelland, and S. Ganguli · 2019
Cited alongside, same era.
High-dimensional dynamics of generalization error in neural networks
M. S. Advani, A. M. Saxe, and H. Sompolinsky · 2020
A unifying view on implicit bias in training linear neural networks
C. Yun, S. Krishnan, and H. Mobahi · 2021
Later among the works it cites.
What happens after SGD reaches zero loss? –a mathematical framework
Z. Li, T. Wang, and S. Arora · 2022
Later among the works it cites.
Scaling ResNets in the large-depth regime
P. Marion, A. Fermanian, G. Biau, and J.-P. Vert · 2022
Later among the works it cites.
Do residual neural networks discretize neural ordinary differential equations?
M. E. Sander, P. Ablin, and G. Peyré · 2022
Later among the works it cites.
Analyzing sharpness along GD trajectory: Progressive sharpening and edge of stability
Z. Wang, Z. Li, and J. Li · 2022
Later among the works it cites.
Stabilize deep ResNet with a sharp scaling factor τ \tau
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process
G. Blanc, N. Gupta, G. Valiant, and P. Valiant · 2020
Cited alongside, same era.
Batch normalization biases residual blocks towards the identity function in deep networks
S. De and S. Smith · 2020
Cited alongside, same era.
Directional convergence and alignment in deep learning
Z. Ji and M. Telgarsky · 2020
Cited alongside, same era.
Fantastic generalization measures and where to find them
Y. Jiang, B. Neyshabur, H. Mobahi, D. Krishnan, and S. Bengio · 2020
Cited alongside, same era.
The large learning rate phase of deep learning: the catapult mechanism
A. Lewkowycz, Y. Bahri, E. Dyer, J. Sohl-Dickstein, and G. Gur-Ari · 2020
Cited alongside, same era.
Unique properties of flat minima in deep networks
R. Mulayoff and T. Michaeli · 2020
Cited alongside, same era.
H. Zhang, D. Yu, M. Yi, W. Chen, and T.-Y. Liu · 2022
Later among the works it cites.
Second-order regression models exhibit progressive sharpening to the edge of stability
A. Agarwala, F. Pedregosa, and J. Pennington · 2023
Later among the works it cites.
A modern look at the relationship between sharpness and generalization
M. Andriushchenko, F. Croce, M. Müller, M. Hein, and N. Flammarion · 2023
Later among the works it cites.
Self-stabilization: The implicit bias of gradient descent at the edge of stability
A. Damian, E. Nichani, and J. D. Lee · 2023
Later among the works it cites.
Same pre-training loss, better downstream: Implicit bias matters for language models
H. Liu, S. M. Xie, Z. Li, and T. Ma · 2023
Later among the works it cites.
On progressive sharpening, flat minima and generalisation
L. E. MacDonald, J. Valmadre, and S. Lucey · 2023
Later among the works it cites.
Fast convergence to non-isolated minima: four equivalent conditions for C 2 \mathrm{C}^{2} functions
Q. Rebjock and N. Boumal · 2023
Later among the works it cites.
Implicit regularization towards rank minimization in ReLU networks
N. Timor, G. Vardi, and O. Shamir · 2023
Later among the works it cites.
On the spectral bias of two-layer linear networks
A. V. Varre, M.-L. Vladarean, L. Pillaud-Vivien, and N. Flammarion · 2023
Later among the works it cites.
The feature speed formula: a flexible approach to scale hyper-parameters of deep neural networks
L. Chizat and P. Netrapalli · 2024
Closest in time.
Implicit regularization of deep residual networks towards neural ODEs
P. Marion, Y.-H. Wu, M. E. Sander, and G. Biau · 2024
Closest in time.
J. Wu, P. L. Bartlett, M. Telgarsky, and B. Yu · 2024
Closest in time.
Tensor programs VI: Feature learning in infinite depth neural networks
G. Yang, D. Yu, C. Zhu, and S. Hayou · 2024
Closest in time.