Fetching the paper…
Reading the bibliography…
In gradient descent dynamics of neural networks, the top eigenvalue of the loss Hessian (sharpness) displays a variety of robust phenomena throughout training.
Chaos in the cubic mapping
Thomas D. Rogers and David C. Whitley · 1983
Earlier work this paper cites.
Chaos in Dynamical Systems
Edward Ott · 2002
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
The mnist database of handwritten digit images for machine learning research
Li Deng · 2012
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, X. Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Earlier work this paper cites.
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Earlier work this paper cites.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms, 2017
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Earlier work this paper cites.
Fantastic generalization measures and where to find them
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 2019
Earlier work this paper cites.
Flax: A neural network library and ecosystem for JAX, 2020
Jonathan Heek, Anselm Levskaya, Avital Oliver, Marvin Ritter, Bertrand Rondepierre, Andreas Steiner, and Marc van Zee · 2020
Earlier work this paper cites.
The break-even point on optimization trajectories of deep neural networks
Stanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit, Jacek Tabor, Kyunghyun Cho*, and Krzysztof Geras* · 2020
Cited alongside, same era.
Stochasticity of deterministic gradient descent: Large learning rate for multiscale objective function
Lingkai Kong and Molei Tao · 2020
Cited alongside, same era.
The large learning rate phase of deep learning: the catapult mechanism
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2020
Cited alongside, same era.
Neural kernels without tangents
Vaishaal Shankar, Alexander W. Fang, Wenshuo Guo, Sara Fridovich-Keil, Ludwig Schmidt, Jonathan Ragan-Kelley, and Benjamin Recht · 2020
Cited alongside, same era.
On the infinite width limit of neural networks with a standard parameterization
Jascha Sohl-Dickstein, Roman Novak, Samuel S. Schoenholz, and Jaehoon Lee · 2020
Cited alongside, same era.
Beyond the edge of stability via two-step gradient updates
Lei Chen and Joan Bruna · 2023
Closest in time.
From stability to chaos: Analyzing gradient descent dynamics in quadratic regression
Xuxing Chen, Krishnakumar Balasubramanian, Promit Ghosal, and Bhavya Agrawalla · 2023
Closest in time.
Self-stabilization: The implicit bias of gradient descent at the edge of stability
Alex Damian, Eshaan Nichani, and Jason D. Lee · 2023
Closest in time.
Critical initialization of wide and deep neural networks through partial jacobians: General theory and applications
Darshil Doshi, Tianyu He, and Andrey Gromov · 2023
Closest in time.
Phase diagram of early training dynamics in deep neural networks: effect of the learning rate, depth, and width
Dayal Singh Kalra and Maissam Barkeshli · 2023
Closest in time.
On the maximum hessian eigenvalue and generalization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Understanding edge-of-stability training dynamics with a minimalist example
Xingyu Zhu, Zixuan Wang, Xiang Wang, Mo Zhou, and Rong Ge · 2020
Cited alongside, same era.
Gradient descent on neural networks typically occurs at the edge of stability
Jeremy Cohen, Simran Kaur, Yuanzhi Li, J Zico Kolter, and Ameet Talwalkar · 2021
Cited alongside, same era.
Tensor programs iv: Feature learning in infinite-width neural networks
Greg Yang and Edward J. Hu · 2021
Cited alongside, same era.
A second-order regression model shows edge of stability behavior
Atish Agarwala, Fabian Pedregosa, and Jeffrey Pennington · 2022
Cited alongside, same era.
Learning threshold neurons via the "edge of stability"
Kwangjun Ahn, Sébastien Bubeck, Sinho Chewi, Yin Tat Lee, Felipe Suarez, and Yi Zhang · 2022
Cited alongside, same era.
Understanding gradient descent on the edge of stability in deep learning
Sanjeev Arora, Zhiyuan Li, and Abhishek Panigrahi · 2022
Cited alongside, same era.
Analyzing sharpness along GD trajectory: Progressive sharpening and edge of stability
Zixuan Wang, Zhouzi Li, and Jian Li · 2022
Cited alongside, same era.
Simran Kaur, Jeremy Cohen, and Zachary Chase Lipton · 2023
Closest in time.
Gradient descent monotonically decreases the sharpness of gradient flow solutions in scalar networks and beyond
Itai Kreisler, Mor Shpigel Nacson, Daniel Soudry, and Yair Carmon · 2023
Closest in time.
On a continuous time model of gradient descent dynamics and instability in deep learning
Mihaela Rosca, Yan Wu, Chongli Qin, and Benoit Dherin · 2023
Closest in time.
Trajectory alignment: Understanding the edge of stability phenomenon via bifurcation theory
Minhak Song and Chulhee Yun · 2023
Closest in time.
Good regularity creates large learning rate implicit biases: edge of stability, balancing, and catapult
Yuqing Wang, Zhenghao Xu, Tuo Zhao, and Molei Tao · 2023
Closest in time.
Implicit bias of gradient descent for logistic regression at the edge of stability
Jingfeng Wu, Vladimir Braverman, and Jason D. Lee · 2023
Closest in time.
Why do learning rates transfer? reconciling optimization and scaling limits for deep learning, 2024
Lorenzo Noci, Alexandru Meterez, Thomas Hofmann, and Antonio Orvieto · 2024
Closest in time.
Beyond the quadratic approximation: The multiscale structure of neural network loss landscapes
ChaoKunin Ma, Lie Wu, and Lexing Ying · 2048
Closest in time.