Fetching the paper…
Reading the bibliography…
Traditional analyses of gradient descent show that when the largest eigenvalue of the Hessian, also known as the sharpness $S(\theta)$, is bounded by $2/\eta$, training is "stable" and the training loss decreases monotonically.
Recursive deep models for semantic compositionality over a sentiment treebank
R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Ng, and C. Potts · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Gradient descent with nonconvex constraints: local concavity determines convergence
R. F. Barber and W. Ha · 2017
Earlier work this paper cites.
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data
G. K. Dziugaite and D. M. Roy · 2017
Earlier work this paper cites.
Three factors influencing minima in sgd
S. Jastrzębski, Z. Kenton, D. Arpit, N. Ballas, A. Fischer, Y. Bengio, and A. J. Storkey · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2017
Earlier work this paper cites.
Exploring generalization in deep learning
B. Neyshabur, S. Bhojanapalli, D. Mcallester, and N. Srebro · 2017
Earlier work this paper cites.
Swish: a self-gated activation function
P. Ramachandran, B. Zoph, and Q. V. Le · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. M. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. VanderPlas, S. Wanderman-Milne, and Q. Zhang · 2018
Earlier work this paper cites.
Don’t decay the learning rate, increase the batch size
S. L. Smith, P.-J. Kindermans, and Q. V. Le · 2018
Earlier work this paper cites.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
L. Wu, C. Ma, and E. Weinan · 2018
Cited alongside, same era.
C. Xing, D. Arpit, C. Tsirigotis, and Y. Bengio · 2018
Cited alongside, same era.
On the relation between the sharpest directions of DNN loss and the SGD step length
S. Jastrzębski, Z. Kenton, N. Ballas, A. Fischer, Y. Bengio, and A. Storkey · 2019
Cited alongside, same era.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Y. Li, C. Wei, and T. Ma · 2019
Cited alongside, same era.
Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process
G. Blanc, N. Gupta, G. Valiant, and P. Valiant · 2020
Cited alongside, same era.
Sharpness-aware minimization for efficiently improving generalization
P. Foret, A. Kleiner, H. Mobahi, and B. Neyshabur · 2021
Later among the works it cites.
A loss curvature perspective on training instability in deep learning, 2021
J. Gilmer, B. Ghorbani, A. Garg, S. Kudugunta, B. Neyshabur, D. Cardoze, G. Dahl, Z. Nado, and O. Firat · 2021
Later among the works it cites.
Second-order regression models exhibit progressive sharpening to the edge of stability, 2022
A. Agarwala, F. Pedregosa, and J. Pennington · 2022
Closest in time.
Understanding the unstable convergence of gradient descent
K. Ahn, J. Zhang, and S. Sra · 2022
Closest in time.
Understanding gradient descent on the edge of stability in deep learning
S. Arora, Z. Li, and A. Panigrahi · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The break-even point on optimization trajectories of deep neural networks
S. Jastrzębski, M. Szymczak, S. Fort, D. Arpit, J. Tabor, K. Cho, and K. Geras · 2020
Cited alongside, same era.
Fantastic generalization measures and where to find them
Y. Jiang, B. Neyshabur, H. Mobahi, D. Krishnan, and S. Bengio · 2020
Cited alongside, same era.
The large learning rate phase of deep learning: the catapult mechanism
A. Lewkowycz, Y. Bahri, E. Dyer, J. Sohl-Dickstein, and G. Gur-Ari · 2020
Cited alongside, same era.
Gradient descent on neural networks typically occurs at the edge of stability
J. Cohen, S. Kaur, Y. Li, J. Z. Kolter, and A. Talwalkar · 2021
Cited alongside, same era.
Label noise SGD provably prefers flat global minimizers
A. Damian, T. Ma, and J. D. Lee · 2021
Cited alongside, same era.
What happens after SGD reaches zero loss? –a mathematical framework
Z. Li, T. Wang, and S. Arora
Cited in the paper.
Analyzing sharpness along gd trajectory: Progressive sharpening and edge of stability
Z. Li, Z. Wang, and J. Li
Cited in the paper.
Closest in time.
On gradient descent convergence beyond the edge of stability
L. Chen and J. Bruna · 2022
Closest in time.
Adaptive gradient methods at the edge of stability, 2022
J. M. Cohen, B. Ghorbani, S. Krishnan, N. Agarwal, S. Medapati, M. Badura, D. Suo, D. Cardoze, Z. Nado, G. E. Dahl, and J. Gilmer · 2022
Closest in time.
Understanding the generalization benefit of normalization layers: Sharpness reduction
K. Lyu, Z. Li, and S. Arora · 2022
Closest in time.
The multiscale structure of neural network loss functions: The effect on optimization and origin
C. Ma, L. Wu, and L. Ying · 2022
Closest in time.
Understanding edge-of-stability training dynamics with a minimalist example, 2022
X. Zhu, Z. Wang, X. Wang, M. Zhou, and R. Ge · 2022
Closest in time.