Fetching the paper…
Reading the bibliography…
The success of deep learning is due, to a large extent, to the remarkable effectiveness of gradient-based optimization methods applied to large neural networks.
“A topological property of real analytic subsets”
Stanislaw Lojasiewicz · 1963
Earlier work this paper cites.
“Gradient methods for minimizing functionals”
Boris Polyak · 1963
Earlier work this paper cites.
“A method for unconstrained convex minimization problem with the rate of convergence O (1/kˆ 2)”
Yurii Nesterov · 1983
Earlier work this paper cites.
“On the local minima free condition of backpropagation learning”
Xiao-Hu Yu and Guo-An Chen · 1995
Earlier work this paper cites.
“Condition numbers of Gaussian random matrices”
Zizhong Chen and Jack Dongarra · 2005
Earlier work this paper cites.
“Numerical optimization”
Jorge Nocedal and Stephen Wright · 2006
Earlier work this paper cites.
“Condition: The geometry of numerical algorithms”
Peter Bürgisser and Felipe Cucker · 2013
Earlier work this paper cites.
“Adam: A method for stochastic optimization”
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
“Deep residual learning for image recognition”
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2016
Earlier work this paper cites.
“On exponential convergence of sgd in non-convex over-parametrized learning”
Raef Bassily, Mikhail Belkin and Siyuan Ma · 2018
Earlier work this paper cites.
“Stability and generalization of learning algorithms that converge to global optima”
Zachary Charles and Dimitris Papailiopoulos · 2018
Earlier work this paper cites.
“The loss landscape of overparameterized neural networks”
Yaim Cooper · 2018
Earlier work this paper cites.
“Gradient descent provably optimizes over-parameterized neural networks”
Simon Du, Xiyu Zhai, Barnabas Poczos and Aarti Singh · 2018
Earlier work this paper cites.
“Neural tangent kernel: Convergence and generalization in neural networks”
Arthur Jacot, Franck Gabriel and Clément Hongler · 2018
Cited alongside, same era.
“Over-parameterized deep neural networks have no strict local minima for any continuous activations”
Dawei Li, Tian Ding and Ruoyu Sun · 2018
Cited alongside, same era.
“A mean field view of the landscape of two-layer neural networks”
Song Mei, Andrea Montanari and Phan-Minh Nguyen · 2018
Cited alongside, same era.
“On the loss landscape of a class of deep neural networks with no bad local valleys”
Quynh Nguyen, Mahesh Mukkamala and Matthias Hein · 2018
Cited alongside, same era.
“Theoretical insights into the optimization landscape of over-parameterized shallow neural networks”
Mahdi Soltanolkotabi, Adel Javanmard and Jason Lee · 2018
Cited alongside, same era.
“Wide neural networks of any depth evolve as linear models under gradient descent”
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein and Jeffrey Pennington · 2019
Later among the works it cites.
“Overparameterized nonlinear learning: Gradient descent takes the shortest path?”
Samet Oymak and Mahdi Soltanolkotabi · 2019
Later among the works it cites.
Samet Oymak and Mahdi Soltanolkotabi · 2019
Later among the works it cites.
“Double descent in the condition number”
Tomaso Poggio, Gil Kur and Andrzej Banburski · 2019
Later among the works it cites.
“A jamming transition from under- to over-parametrization affects generalization in deep learning”
S Spigler, M Geiger, S d’Ascoli, L Sagun, G Biroli and M Wyart · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Stochastic gradient descent optimizes over-parameterized deep relu networks”
Difan Zou, Yuan Cao, Dongruo Zhou and Quanquan Gu · 2018
Cited alongside, same era.
“A Convergence Theory for Deep Learning via Over-Parameterization”
Zeyuan Allen-Zhu, Yuanzhi Li and Zhao Song · 2019
Cited alongside, same era.
“Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks”
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li and Ruosong Wang · 2019
Cited alongside, same era.
“Gradient descent with identity initialization efficiently learns positive-definite linear transformations by deep residual networks”
Peter Bartlett, David Helmbold and Philip Long · 2019
Cited alongside, same era.
“Reconciling modern machine-learning practice and the classical bias–variance trade-off”
Mikhail Belkin, Daniel Hsu, Siyuan Ma and Soumik Mandal · 2019
Cited alongside, same era.
“Gradient Descent Finds Global Minima of Deep Neural Networks”
Simon Du, Jason Lee, Haochuan Li, Liwei Wang and Xiyu Zhai · 2019
Cited alongside, same era.
“Path length bounds for gradient descent and flow”
Chirag Gupta, Sivaraman Balakrishnan and Aaditya Ramdas · 2019
Cited alongside, same era.
Later among the works it cites.
“Fast and Faster Convergence of SGD for Over-Parameterized Models and an Accelerated Perceptron”
Sharan Vaswani, Francis Bach and Mark Schmidt · 2019
Later among the works it cites.
“Language models are few-shot learners”
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry and Amanda Askell · 2020
Closest in time.
“No Spurious Local Minima: on the Optimization Landscapes of Wide and Deep Neural Networks”
Johannes Lederer · 2020
Closest in time.
“Accelerating SGD with momentum for over-parameterized learning”
Chaoyue Liu and Mikhail Belkin · 2020
Closest in time.
“On the linearity of large non-linear models: when and why the tangent kernel is constant”
Chaoyue Liu, Libin Zhu and Mikhail Belkin · 2020
Closest in time.
“Beyond convexity—Contraction and global convergence of gradient descent”
Patrick Wensing and Jean-Jacques Slotine · 2020
Closest in time.
“Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity”
William Fedus, Barret Zoph and Noam Shazeer · 2021
Closest in time.