Fetching the paper…
Reading the bibliography…
In this work, we investigate the mechanism underlying loss spikes observed during neural network training.
Reflections after refereeing papers for nips
Leo Breiman · 1995
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Lanczos algorithms for large symmetric eigenvalue computations: Vol. I: Theory
Jane K Cullum and Ralph A Willoughby · 2002
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Earlier work this paper cites.
Three factors influencing minima in sgd
Stanislaw Jastrzebski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Earlier work this paper cites.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2017
Earlier work this paper cites.
Towards understanding generalization of deep learning: Perspective of loss landscapes
Lei Wu, Zhanxing Zhu, et al · 2017
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
P Izmailov, AG Wilson, D Podoprikhin, D Vetrov, and T Garipov · 2018
Earlier work this paper cites.
Don’t use large mini-batches, use local sgd
Tao Lin, Sebastian U Stich, Kumar Kshitij Patel, and Martin Jaggi · 2018
Earlier work this paper cites.
Gradient descent quantizes relu network features
Hartmut Maennel, Olivier Bousquet, and Sylvain Gelly · 2018
Earlier work this paper cites.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
Lei Wu, Chao Ma, and Weinan E · 2018
Earlier work this paper cites.
Chen Xing, Devansh Arpit, Christos Tsirigotis, and Yoshua Bengio · 2018
Earlier work this paper cites.
Frequency-aware reconstruction of fluid simulations with generative networks
Simon Biland, Vinicius C Azevedo, Byungsoo Kim, and Barbara Solenthaler · 2019
Earlier work this paper cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2019
Earlier work this paper cites.
Gradient descent finds global minima of deep neural networks
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Earlier work this paper cites.
Semi-flat minima and saddle points by embedding neural networks to overparameterization
Kenji Fukumizu, Shoichiro Yamaguchi, Yoh-ichi Mototake, and Mirai Tanaka · 2019
Cited alongside, same era.
Control batch size and learning rate to generalize well: Theoretical and empirical evidence
Fengxiang He, Tongliang Liu, and Dacheng Tao · 2019
Cited alongside, same era.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Yuanzhi Li, Colin Wei, and Tengyu Ma · 2019
Cited alongside, same era.
On the spectral bias of deep neural networks
Nasim Rahaman, Devansh Arpit, Aristide Baratin, Felix Draxler, Min Lin, Fred A Hamprecht, Yoshua Bengio, and Aaron Courville · 2019
Cited alongside, same era.
The convergence rate of neural networks for learned functions of different frequencies
Basri Ronen, David Jacobs, Yoni Kasten, and Shira Kritchman · 2019
Cited alongside, same era.
Towards understanding the condensation of neural networks at initial training
Hanxu Zhou, Qixuan Zhou, Tao Luo, Yaoyu Zhang, and Zhi-Qin John Xu · 2021
Later among the works it cites.
Second-order regression models exhibit progressive sharpening to the edge of stability
Atish Agarwala, Fabian Pedregosa, and Jeffrey Pennington · 2022
Later among the works it cites.
Understanding the unstable convergence of gradient descent
Kwangjun Ahn, Jingzhao Zhang, and Suvrit Sra · 2022
Later among the works it cites.
Sgd with large step sizes learns sparse features
Maksym Andriushchenko, Aditya Varre, Loucas Pillaud-Vivien, and Nicolas Flammarion · 2022
Later among the works it cites.
Understanding gradient descent on the edge of stability in deep learning
Sanjeev Arora, Zhiyuan Li, and Abhishek Panigrahi · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Training behavior of deep neural network in frequency domain
Zhi-Qin John Xu, Yaoyu Zhang, and Yanyang Xiao · 2019
Cited alongside, same era.
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2020
Cited alongside, same era.
Adaptive activation functions accelerate convergence in deep and physics-informed neural networks
Ameya D Jagtap, Kenji Kawaguchi, and George Em Karniadakis · 2020
Cited alongside, same era.
The large learning rate phase of deep learning: the catapult mechanism
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2020
Cited alongside, same era.
Multi-scale deep neural network (mscalednn) for solving poisson-boltzmann equation in complex domains
Ziqi Liu, Wei Cai, and Zhi-Qin John Xu · 2020
Cited alongside, same era.
An analytic theory of shallow networks dynamics for hinge loss classification
Franco Pellegrini and Giulio Biroli · 2020
Cited alongside, same era.
Frequency principle: Fourier analysis sheds light on deep neural networks
Zhi-Qin John Xu, Yaoyu Zhang, Tao Luo, Yanyang Xiao, and Zheng Ma · 2020
Cited alongside, same era.
Later among the works it cites.
On gradient descent convergence beyond the edge of stability
Lei Chen and Joan Bruna · 2022
Later among the works it cites.
Self-stabilization: The implicit bias of gradient descent at the edge of stability
Alex Damian, Eshaan Nichani, and Jason D Lee · 2022
Later among the works it cites.
Understanding the generalization benefit of normalization layers: Sharpness reduction
Kaifeng Lyu, Zhiyuan Li, and Sanjeev Arora · 2022
Later among the works it cites.
Analyzing sharpness along gd trajectory: Progressive sharpening and edge of stability
Zixuan Wang, Zhouzi Li, and Jian Li · 2022
Later among the works it cites.
Overview frequency principle/spectral bias in deep learning
Zhi-Qin John Xu, Yaoyu Zhang, and Tao Luo · 2022
Later among the works it cites.
Embedding principle: a hierarchical structure of loss landscape of deep neural networks
Yaoyu Zhang, Yuqing Li, Zhongwang Zhang, Tao Luo, and Zhi-Qin John Xu · 2022
Later among the works it cites.
Empirical phase diagram for three-layer neural networks with infinite width
Hanxu Zhou, Qixuan Zhou, Zhenyuan Jin, Tao Luo, Yaoyu Zhang, and Zhi-Qin John Xu · 2022
Later among the works it cites.
Understanding edge-of-stability training dynamics with a minimalist example
Xingyu Zhu, Zixuan Wang, Xiang Wang, Mo Zhou, and Rong Ge · 2022
Later among the works it cites.
Phase diagram of initial condensation for two-layer neural networks
Zhengan Chen, Yuqing Li, Tao Luo, Zhangchen Zhou, and Zhi-Qin John Xu · 2023
Closest in time.
Flat minima generalize for low-rank matrix recovery
Lijun Ding, Dmitriy Drusvyatskiy, Maryam Fazel, and Zaid Harchaoui · 2024
Closest in time.
Understanding the generalization benefits of late learning rate decay
Yinuo Ren, Chao Ma, and Lexing Ying · 2024
Closest in time.
Beyond the quadratic approximation: The multiscale structure of neural network loss landscapes
Chao Ma, Daniel Kunin, Lei Wu, and Lexing Ying · 2048
Closest in time.