Fetching the paper…
Reading the bibliography…
In this paper, we first present an explanation regarding the common occurrence of spikes in the training loss when neural networks are trained with stochastic gradient descent (SGD).
“A stochastic approximation method”
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
“A method for unconstrained convex minimization problem with the rate of convergence O (1/kˆ 2)”
Yurii Nesterov · 1983
Earlier work this paper cites.
“Investigating smooth multiple regression by the method of average derivatives”
Wolfgang Härdle and Thomas Stoker · 1989
Earlier work this paper cites.
“Simplifying neural nets by discovering flat minima”
Sepp Hochreiter and Jürgen Schmidhuber · 1994
Earlier work this paper cites.
“A database for handwritten text recognition research”
J.. Hull · 1994
Earlier work this paper cites.
“Flat minima”
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
“On the momentum term in gradient descent learning algorithms”
Ning Qian · 1999
Earlier work this paper cites.
“Structure adaptive approach for dimension reduction”
Marian Hristache, Anatoli Juditsky, Jorg Polzehl and Vladimir Spokoiny · 2001
Earlier work this paper cites.
“Efficient backprop”
Yann LeCun, Léon Bottou, Genevieve Orr and Klaus-Robert Müller · 2002
Earlier work this paper cites.
“An adaptive estimation of dimension reduction space”
Yingcun Xia, Howell Tong, Wai Li and Li-Xing Zhu · 2002
Earlier work this paper cites.
“Learning multiple layers of features from tiny images”
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
“Adaptive subgradient methods for online learning and stochastic optimization.”
John Duchi, Elad Hazan and Yoram Singer · 2011
Earlier work this paper cites.
“Reading digits in natural images with unsupervised feature learning”, 2011
Yuval Netzer, Tao Wang, Adam Coates, Alessandro Bissacco, Bo Wu and Andrew Ng · 2011
Earlier work this paper cites.
“Adadelta: an adaptive learning rate method”
Matthew Zeiler · 2012
Earlier work this paper cites.
URL: http://www.cs.toronto.edu/~tijmen/csc321/slides/lecture_slides_lec6.pdf
Geoffrey Hinton, 2014 · 2014
Earlier work this paper cites.
“A consistent estimator of the expected gradient outerproduct”
Shubhendu Trivedi, Jialei Wang, Samory Kpotufe and Gregory Shakhnarovich · 2014
Earlier work this paper cites.
“Adam: A Method for Stochastic Optimization”
Diederik Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
“Deep learning”
Yann LeCun, Yoshua Bengio and Geoffrey Hinton · 2015
Earlier work this paper cites.
“Deep Learning Face Attributes in the Wild”
Ziwei Liu, Ping Luo, Xiaogang Wang and Xiaoou Tang · 2015
Earlier work this paper cites.
“Deep residual learning for image recognition”
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2016
Earlier work this paper cites.
“An overview of gradient descent optimization algorithms”
Sebastian Ruder · 2016
Earlier work this paper cites.
“Wide Residual Networks”
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
“Sharp minima can generalize for deep nets”
Laurent Dinh, Razvan Pascanu, Samy Bengio and Yoshua Bengio · 2017
Earlier work this paper cites.
“Accurate, large minibatch sgd: Training imagenet in 1 hour”
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia and Kaiming He · 2017
Earlier work this paper cites.
“Densely connected convolutional networks”
Gao Huang, Zhuang Liu, Laurens Van and Kilian Weinberger · 2017
Earlier work this paper cites.
“Three factors influencing minima in sgd”
Stanisław Jastrzębski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio and Amos Storkey · 2017
Earlier work this paper cites.
“On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima”
Nitish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy and Ping Tang · 2017
Earlier work this paper cites.
“Improving generalization performance by switching from adam to sgd”
Nitish Keskar and Richard Socher · 2017
Earlier work this paper cites.
“Exploring generalization in deep learning”
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester and Nati Srebro · 2017
Earlier work this paper cites.
“Cyclical learning rates for training neural networks”
Leslie Smith · 2017
Cited alongside, same era.
“Towards understanding generalization of deep learning: Perspective of loss landscapes”
Lei Wu and Zhanxing Zhu · 2017
Cited alongside, same era.
“Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms”
Han Xiao, Kashif Rasul and Roland Vollgraf · 2017
Cited alongside, same era.
“Averaging weights leads to wider optima and better generalization”
P Izmailov, AG Wilson, D Podoprikhin, D Vetrov and T Garipov · 2018
Cited alongside, same era.
“Neural tangent kernel: Convergence and generalization in neural networks”
Arthur Jacot, Franck Gabriel and Clément Hongler · 2018
Cited alongside, same era.
“What can linearized neural networks actually say about generalization?”
Guillermo Ortiz-Jiménez, Seyed-Mohsen Moosavi-Dezfooli and Pascal Frossard · 2021
Later among the works it cites.
“On the Origin of Implicit Regularization in Stochastic Gradient Descent”
Samuel Smith, Benoit Dherin, David Barrett and Soham De · 2021
Later among the works it cites.
“A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima”
Zeke Xie, Issei Sato and Masashi Sugiyama · 2021
Later among the works it cites.
“Tensor programs iv: Feature learning in infinite-width neural networks”
Greg Yang and Edward Hu · 2021
Later among the works it cites.
“Understanding the unstable convergence of gradient descent”
Kwangjun Ahn, Jingzhao Zhang and Suvrit Sra · 2022
Later among the works it cites.
“Understanding gradient descent on the edge of stability in deep learning”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“An alternative view: When does SGD escape local minima?”
Bobby Kleinberg, Yuanzhi Li and Yang Yuan · 2018
Cited alongside, same era.
“Revisiting small batch training for deep neural networks”
Dominic Masters and Carlo Luschi · 2018
Cited alongside, same era.
“Myrtle Network”, https://myrtle.ai/ , 2018
Myrtle.ai · 2018
Cited alongside, same era.
Chen Xing, Devansh Arpit, Christos Tsirigotis and Yoshua Bengio · 2018
Cited alongside, same era.
“Gradient descent finds global minima of deep neural networks”
Simon Du, Jason Lee, Haochuan Li, Liwei Wang and Xiyu Zhai · 2019
Cited alongside, same era.
“On the Relation Between the Sharpest Directions of DNN Loss and the SGD Step Length”
Stanisław Jastrzębski, Zachary Kenton, Nicolas Ballas, Asja Fischer, Yoshua Bengio and Amost Storkey · 2019
Cited alongside, same era.
“Wide neural networks of any depth evolve as linear models under gradient descent”
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein and Jeffrey Pennington · 2019
Cited alongside, same era.
Sanjeev Arora, Zhiyuan Li and Abhishek Panigrahi · 2022
Later among the works it cites.
“Neural Networks as Kernel Learners: The Silent Alignment Effect”
Alexander Atanasov, Blake Bordelon and Cengiz Pehlevan · 2022
Later among the works it cites.
“Stochastic Training is Not Necessary for Generalization”
Jonas Geiping, Micah Goldblum, Phil Pope, Michael Moeller and Tom Goldstein · 2022
Later among the works it cites.
“Loss landscapes and optimization in over-parameterized non-linear systems and neural networks”
Chaoyue Liu, Libin Zhu and Mikhail Belkin · 2022
Later among the works it cites.
“Transition to Linearity of Wide Neural Networks is an Emerging Property of Assembling Weak Models”
Chaoyue Liu, Libin Zhu and Misha Belkin · 2022
Later among the works it cites.
“Evolution of neural tangent kernels under benign and adversarial training”
Noel Loo, Ramin Hasani, Alexander Amini and Daniela Rus · 2022
Later among the works it cites.
“Large Learning Rate Tames Homogeneity: Convergence and Balancing Effect”
Yuqing Wang, Minshuo Chen, Tuo Zhao and Molei Tao · 2022
Later among the works it cites.
“Analyzing sharpness along gd trajectory: Progressive sharpening and edge of stability”
Zixuan Wang, Zhouzi Li and Jian Li · 2022
Later among the works it cites.
“Transition to linearity of general neural networks with directed acyclic graph architecture”
Libin Zhu, Chaoyue Liu and Misha Belkin · 2022
Later among the works it cites.
“SAM operates far from home: eigenvalue regularization as a dynamical phenomenon”
Atish Agarwala and Yann Dauphin · 2023
Closest in time.
“Second-order regression models exhibit progressive sharpening to the edge of stability”
Atish Agarwala, Fabian Pedregosa and Jeffrey Pennington · 2023
Closest in time.
“Neural tangent kernel at initialization: linear width suffices”
Arindam Banerjee, Pedro Cisneros-Velarde, Libin Zhu and Mikhail Belkin · 2023
Closest in time.
“Mechanism of feature learning in convolutional neural networks”
Daniel Beaglehole, Adityanarayanan Radhakrishnan, Parthe Pandit and Mikhail Belkin · 2023
Closest in time.
“Self-Stabilization: The Implicit Bias of Gradient Descent at the Edge of Stability”
Alex Damian, Eshaan Nichani and Jason. Lee · 2023
Closest in time.
“Phase diagram of early training dynamics in deep neural networks: effect of the learning rate, depth, and width”
Dayal Kalra and Maissam Barkeshli · 2023
Closest in time.
“Catapult Dynamics and Phase Transitions in Quadratic Nets”
David Meltzer and Junyu Liu · 2023
Closest in time.
“Efficient Estimation of the Central Mean Subspace via Smoothed Gradient Outer Products”
Gan Yuan, Mingyue Xu, Samory Kpotufe and Daniel Hsu · 2023
Closest in time.
“Loss Spike in Training Neural Networks”
Zhongwang Zhang and Zhi-Qin Xu · 2023
Closest in time.
“Flat minima generalize for low-rank matrix recovery”
Lijun Ding, Dmitriy Drusvyatskiy, Maryam Fazel and Zaid Harchaoui · 2024
Closest in time.
“Mechanism for feature learning in neural networks and backpropagation-free machine learning models”
Adityanarayanan Radhakrishnan, Daniel Beaglehole, Parthe Pandit and Mikhail Belkin · 2024
Closest in time.
“Linear Recursive Feature Machines provably recover low-rank matrices”
Adityanarayanan Radhakrishnan, Mikhail Belkin and Dmitriy Drusvyatskiy · 2024
Closest in time.
“Quadratic models for understanding catapult dynamics of neural networks”
Libin Zhu, Chaoyue Liu, Adityanarayanan Radhakrishnan and Mikhail Belkin · 2024
Closest in time.
“An improved analysis of training over-parameterized deep neural networks”
Difan Zou and Quanquan Gu · 2062
Closest in time.