Fetching the paper…
Reading the bibliography…
We identify a new phenomenon in neural network optimization which arises from the interaction of depth and a particular heavy-tailed structure in natural data.
Learning Multiple Layers of Features from Tiny Images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Eigenvalues of the hessian in deep learning: Singularity and beyond
Levent Sagun, Leon Bottou, and Yann LeCun · 2016
Earlier work this paper cites.
Analysis of deep neural networks with extended data jacobian matrix
Shengjie Wang, Abdel-rahman Mohamed, Rich Caruana, Jeff Bilmes, Matthai Plilipose, Matthew Richardson, Krzysztof Geras, Gregor Urban, and Ozlem Aslan · 2016
Earlier work this paper cites.
Three factors influencing minima in sgd
Stanisław Jastrzȩbski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Earlier work this paper cites.
Empirical analysis of the hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V Ugur Guney, Yann Dauphin, and Leon Bottou · 2017
Earlier work this paper cites.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
Simon S Du, Wei Hu, and Jason D Lee · 2018
Earlier work this paper cites.
The full spectrum of deepnet hessians at scale: Dynamics with sgd training and sample size
Vardan Papyan · 2018
Earlier work this paper cites.
How does batch normalization help optimization?
Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, and Aleksander Madry · 2018
Earlier work this paper cites.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
Lei Wu, Chao Ma, and Weinan E · 2018
Earlier work this paper cites.
Chen Xing, Devansh Arpit, Christos Tsirigotis, and Yoshua Bengio · 2018
Earlier work this paper cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Earlier work this paper cites.
Emergent properties of the local geometry of neural loss landscapes
Stanislav Fort and Surya Ganguli · 2019
Earlier work this paper cites.
An investigation into neural net optimization via hessian eigenvalue density
Behrooz Ghorbani, Shankar Krishnan, and Ying Xiao · 2019
Earlier work this paper cites.
Openwebtext corpus, 2019
Aaron Gokaslan, Vanya Cohen, Ellie Pavlick, and Stefanie Tellex · 2019
Earlier work this paper cites.
On the relation between the sharpest directions of DNN loss and the SGD step length
Stanisław Jastrzȩbski, Zachary Kenton, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amost Storkey · 2019
Earlier work this paper cites.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Yuanzhi Li, Colin Wei, and Tengyu Ma · 2019
Earlier work this paper cites.
Do deep neural networks learn shallow learnable examples first?
Karttikeya Mangalam and Vinay Uday Prabhu · 2019
Earlier work this paper cites.
Sgd on neural networks learns functions of increasing complexity
Preetum Nakkiran, Gal Kaplun, Dimitris Kalimeris, Tristan Yang, Benjamin L Edelman, Fred Zhang, and Boaz Barak · 2019
Earlier work this paper cites.
Generalization guarantees for neural networks via harnessing the low-rank structure of the jacobian
Samet Oymak, Zalan Fabian, Mingchen Li, and Mahdi Soltanolkotabi · 2019
Cited alongside, same era.
Measurements of three-level hierarchical structure in the outliers in the spectrum of deepnet hessians
Vardan Papyan · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Deep learning generalizes because the parameter-function map is biased towards simple functions
Guillermo Valle-Perez, Chico Q. Camargo, and Ard A. Louis · 2019
Cited alongside, same era.
The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects
Zhanxing Zhu, Jingfeng Wu, Bing Yu, Lei Wu, and Jinwen Ma · 2019
Cited alongside, same era.
Hidden progress in deep learning: SGD learns parities near the computational limit
Boaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham M. Kakade, eran malach, and Cyril Zhang · 2022
Later among the works it cites.
Gradient descent on neurons and its link to approximate second-order optimization
Frederik Benzing · 2022
Later among the works it cites.
Adaptive gradient methods at the edge of stability
Jeremy M Cohen, Behrooz Ghorbani, Shankar Krishnan, Naman Agarwal, Sourabh Medapati, Michal Badura, Daniel Suo, David Cardoze, Zachary Nado, George E Dahl, et al · 2022
Later among the works it cites.
Self-stabilization: The implicit bias of gradient descent at the edge of stability
Alex Damian, Eshaan Nichani, and Jason D. Lee · 2022
Later among the works it cites.
How to fine-tune vision models with sgd
Ananya Kumar, Ruoqi Shen, Sébastien Bubeck, and Suriya Gunasekar · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
The break-even point on optimization trajectories of deep neural networks
Stanisław Jastrzȩbski, Maciej Szymczak, Stanislav Fort, Devansh Arpit, Jacek Tabor, Kyunghyun Cho*, and Krzysztof Geras* · 2020
Cited alongside, same era.
Neural spectrum alignment: Empirical study
Dmitry Kopitkov and Vadim Indelman · 2020
Cited alongside, same era.
The large learning rate phase of deep learning: the catapult mechanism
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2020
Cited alongside, same era.
Hessian based analysis of sgd for deep nets: Dynamics and generalization
Xinyan Li, Qilong Gu, Yingxue Zhou, Tiancong Chen, and Arindam Banerjee · 2020
Cited alongside, same era.
Unique properties of flat minima in deep networks
Rotem Mulayoff and Tomer Michaeli · 2020
Cited alongside, same era.
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2020
Cited alongside, same era.
Later among the works it cites.
Beyond the quadratic approximation: the multiscale structure of neural network loss landscapes
Chao Ma, Daniel Kunin, Lei Wu, and Lexing Ying · 2022
Later among the works it cites.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra · 2022
Later among the works it cites.
Elan Rosenfeld, Pradeep Ravikumar, and Andrej Risteski · 2022
Later among the works it cites.
The slingshot mechanism: An empirical study of adaptive optimizers and the grokking phenomenon
Vimal Thilak, Etai Littwin, Shuangfei Zhai, Omid Saremi, Roni Paiss, and Joshua Susskind · 2022
Later among the works it cites.
Analyzing sharpness along GD trajectory: Progressive sharpening and edge of stability
Zixuan Wang, Zhouzi Li, and Jian Li · 2022
Later among the works it cites.
The alignment property of SGD noise and how it helps select flat minima: A stability analysis
Lei Wu, Mingze Wang, and Weijie J Su · 2022
Later among the works it cites.
Symbolic discovery of optimization algorithms
Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Kaiyuan Wang, Yao Liu, Hieu Pham, Xuanyi Dong, Thang Luong, Cho-Jui Hsieh, et al · 2023
Closest in time.
On the lipschitz constant of deep networks and double descent
Matteo Gamba, Hossein Azizpour, and Mårten Björkman · 2023
Closest in time.
Gradient descent monotonically decreases the sharpness of gradient flow solutions in scalar networks and beyond
Itai Kreisler, Mor Shpigel Nacson, Daniel Soudry, and Yair Carmon · 2023
Closest in time.
Noise is not the main factor behind the gap between sgd and adam on transformers, but sign descent might be
Frederik Kunstner, Jacques Chen, Jonathan Wilder Lavington, and Mark Schmidt · 2023
Closest in time.
On progressive sharpening, flat minima and generalisation
Lachlan Ewen MacDonald, Jack Valmadre, and Simon Lucey · 2023
Closest in time.
A tale of two circuits: Grokking as competition of sparse and dense subnetworks
William Merrill, Nikolaos Tsilivis, and Aman Shukla · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt · 2023
Closest in time.
Predicting grokking long before it happens: A look into the loss landscape of models which grok
Pascal Notsawo Jr, Hattie Zhou, Mohammad Pezeshki, Irina Rish, Guillaume Dumas, et al · 2023
Closest in time.
Toward understanding why adam converges faster than sgd for transformers
Yan Pan and Yuanzhi Li · 2023
Closest in time.
Sharpness minimization algorithms do not only minimize sharpness to achieve better generalization
Kaiyue Wen, Tengyu Ma, and Zhiyuan Li · 2023
Closest in time.
Understanding edge-of-stability training dynamics with a minimalist example
Xingyu Zhu, Zixuan Wang, Xiang Wang, Mo Zhou, and Rong Ge · 2023
Closest in time.