Fetching the paper…
Reading the bibliography…
Recent empirical and theoretical work has shown that the dynamics of the large eigenvalues of the training loss Hessian have some remarkably robust features across models and datasets in the full batch regime.
An image is worth 16x16 words: Transformers for image recognition at scale, 2021
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2010
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Don’t Decay the Learning Rate, Increase the Batch Size
Samuel L. Smith, Pieter-Jan Kindermans, and Quoc V. Le · 2017
Earlier work this paper cites.
Three Factors Influencing Minima in SGD, September 2018
Stanisław Jastrzkebski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2018
Earlier work this paper cites.
Neural Tangent Kernel: Convergence and Generalization in Neural Networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Earlier work this paper cites.
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour, April 2018
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2018
Earlier work this paper cites.
An Investigation into Neural Net Optimization via Hessian Eigenvalue Density
Behrooz Ghorbani, Shankar Krishnan, and Ying Xiao · 2019
Earlier work this paper cites.
Wide Neural Networks of Any Depth Evolve as Linear Models Under Gradient Descent
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Earlier work this paper cites.
Mean-field theory of two-layers neural networks: Dimension-free bounds and kernel limit
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Earlier work this paper cites.
Measuring the Effects of Data Parallelism on Neural Network Training
Christopher J. Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E. Dahl · 2019
Earlier work this paper cites.
The Break-Even Point on Optimization Trajectories of Deep Neural Networks, February 2020
Stanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit, Jacek Tabor, Kyunghyun Cho, and Krzysztof Geras · 2020
Earlier work this paper cites.
The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of Generalization
Ben Adlam and Jeffrey Pennington · 2020
Cited alongside, same era.
SGD in the Large: Average-case Analysis, Asymptotics, and Stepsize Criticality
Courtney Paquette, Kiwon Lee, Fabian Pedregosa, and Elliot Paquette · 2021
Cited alongside, same era.
On Linear Stability of SGD and Input-Smoothness of Neural Networks
Chao Ma and Lexing Ying · 2021
Cited alongside, same era.
MLP-Mixer: An all-MLP Architecture for Vision, June 2021
Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, and Alexey Dosovitskiy · 2021
Cited alongside, same era.
A Loss Curvature Perspective on Training Instabilities of Deep Learning Models
Justin Gilmer, Behrooz Ghorbani, Ankush Garg, Sneha Kudugunta, Behnam Neyshabur, David Cardoze, George Edward Dahl, Zachary Nado, and Orhan Firat · 2022
Cited alongside, same era.
Sharpness-aware Minimization for Efficiently Improving Generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2022
Later among the works it cites.
The Implicit Regularization of Dynamical Stability in Stochastic Gradient Descent
Lei Wu and Weijie J. Su · 2023
Later among the works it cites.
Exact Mean Square Linear Stability Analysis for SGD, June 2023
Rotem Mulayoff and Tomer Michaeli · 2023
Later among the works it cites.
From high-dimensional & mean-field dynamics to dimensionless ODEs: A unifying approach to SGD in two-layers networks, February 2023
Luca Arnaboldi, Ludovic Stephan, Florent Krzakala, and Bruno Loureiro · 2023
Later among the works it cites.
How Does Sharpness-Aware Minimization Minimize Sharpness?, January 2023
Kaiyue Wen, Tengyu Ma, and Zhiyuan Li · 2023
Later among the works it cites.
Hessian Inertia in Neural Networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Self-Stabilization: The Implicit Bias of Gradient Descent at the Edge of Stability, September 2022
Alex Damian, Eshaan Nichani, and Jason D. Lee · 2022
Cited alongside, same era.
Second-order regression models exhibit progressive sharpening to the edge of stability, October 2022
Atish Agarwala, Fabian Pedregosa, and Jeffrey Pennington · 2022
Cited alongside, same era.
The alignment property of SGD noise and how it helps select flat minima: A stability analysis, October 2022
Lei Wu, Mingze Wang, and Weijie Su · 2022
Cited alongside, same era.
High-dimensional limit theorems for SGD: Effective dynamics and critical scaling
Gerard Ben Arous, Reza Gheissari, and Aukosh Jagannath · 2022
Cited alongside, same era.
Quadratic models for understanding neural network dynamics, May 2022
Libin Zhu, Chaoyue Liu, Adityanarayanan Radhakrishnan, and Mikhail Belkin · 2022
Cited alongside, same era.
Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability
Jeremy Cohen, Simran Kaur, Yuanzhi Li, J. Zico Kolter, and Ameet Talwalkar
Cited in the paper.
Adaptive Gradient Methods at the Edge of Stability, July 2022b
Jeremy M. Cohen, Behrooz Ghorbani, Shankar Krishnan, Naman Agarwal, Sourabh Medapati, Michal Badura, Daniel Suo, David Cardoze, Zachary Nado, George E. Dahl, and Justin Gilmer
Cited in the paper.
Xuchan Bao, Alberto Bietti, Aaron Defazio, and Vivien Cabannes · 2023
Later among the works it cites.
Hitting the High-Dimensional Notes: An ODE for SGD learning dynamics on GLMs and multi-index models, August 2023
Elizabeth Collins-Woodfin, Courtney Paquette, Elliot Paquette, and Inbar Seroussi · 2023
Later among the works it cites.
On the Interplay Between Stepsize Tuning and Progressive Sharpening, December 2023
Vincent Roulet, Atish Agarwala, and Fabian Pedregosa · 2023
Later among the works it cites.
SAM operates far from home: Eigenvalue regularization as a dynamical phenomenon
Atish Agarwala and Yann Dauphin · 2023
Later among the works it cites.
Neglected Hessian component explains mysteries in Sharpness regularization, January 2024
Yann N. Dauphin, Atish Agarwala, and Hossein Mobahi · 2024
Closest in time.