Fetching the paper…
Reading the bibliography…
Although the optimization objectives for learning neural networks are highly non-convex, gradient-based methods have been wildly successful at learning neural networks in practice.
“A topological property of real analytic subsets”
S. Lojasiewicz · 1963
Earlier work this paper cites.
“Gradient methods for the minimisation of functionals”
B.. Polyak · 1963
Earlier work this paper cites.
“On sufficiency of the Kuhn-Tucker conditions”
Morgan Hanson · 1981
Earlier work this paper cites.
“Invex functions and duality”
B.. Craven and B.. Glover · 1985
Earlier work this paper cites.
“Error bounds and convergence analysis of feasible descent methods: a general approach”
Zhi Quan Luo and Paul Tseng · 1993
Earlier work this paper cites.
“Gradient methods for convex minimization: better rates under weaker conditions”
Hui Zhang and Wotao Yin · 2013
Earlier work this paper cites.
“Escaping From Saddle Points — Online Stochastic Gradient for Tensor Decomposition”
Rong Ge, Furong Huang, Chi Jin and Yang Yuan · 2015
Earlier work this paper cites.
“An Asynchronous Parallel Stochastic Coordinate Descent Algorithm”
Ji Liu, Stephen. Wright, Christopher Re, Victor Bittorf and Srikrishna Sridhar · 2015
Earlier work this paper cites.
“Variance reduction for faster non-convex optimization”
Zeyuan Allen-Zhu and Elad Hazan · 2016
Earlier work this paper cites.
“Linear Convergence of Gradient and Proximal-Gradient Methods Under the Polyak-Lojasiewicz Condition”
Hamed Karimi, Julie Nutini and Mark Schmidt · 2016
Earlier work this paper cites.
“Stochastic Variance Reduction for Nonconvex Optimization”
Sashank. Reddi, Ahmed Hefny, Suvrit Sra, Barnabas Poczos and Alex Smola · 2016
Earlier work this paper cites.
“Identity Matters in Deep Learning”
Moritz Hardt and Tengyu Ma · 2017
Earlier work this paper cites.
“Stochastic recursive gradient algorithm for nonconvex optimization”
Lam Nguyen, Jie Liu, Katya Scheinberg and Martin Takac · 2017
Earlier work this paper cites.
“Diverse Neural Network Learns True Target Functions”
Bo Xie, Yingyu Liang and Le Song · 2017
Cited alongside, same era.
“Characterization of Gradient Dominance and Regularity Conditions for Neural Networks”
Yi Zhou and Yingbin Liang · 2017
Cited alongside, same era.
“Natasha 2: Faster Non-Convex Optimization Than SGD”
Zeyuan Allen-Zhu · 2018
Cited alongside, same era.
“SGD Learns Over-parameterized Networks that Provably Generalize on Linearly Separable Data”
Alon Brutzkus, Amir Globerson, Eran Malach and Shai Shalev-Shwartz · 2018
Cited alongside, same era.
“Stability and Generalization of Learning Algorithms that Converge to Global Optima”
Zachary Charles and Dimitris Papailiopoulos · 2018
Cited alongside, same era.
“On the Global Convergence of Gradient Descent for Over-parameterized Models using Optimal Transport”
“Time/Accuracy Tradeoffs for Learning a ReLU with respect to Gaussian Marginals”
Surbhi Goel, Sushrut Karmalkar and Adam. Klivans · 2019
Later among the works it cites.
“Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit”
Song Mei, Theodor Misiakiewicz and Andrea Montanari · 2019
Later among the works it cites.
“Linear Convergence of First Order Methods for Non-Strongly Convex Optimization”
I. Necoara, Yu. Nesterov and F. Glineur · 2019
Later among the works it cites.
“Exponential Convergence Time of Gradient Descent for One-Dimensional Deep Linear Neural Networks”
Ohad Shamir · 2019
Later among the works it cites.
“Gradient descent optimizes over-parameterized deep ReLU networks”
Difan Zou, Yuan Cao, Dongruo Zhou and Quanquan Gu · 2019
Later among the works it cites.
“Generalization Error Bounds of Gradient Descent for Learning Over-parameterized Deep ReLU Networks”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lenaic Chizat and Francis Bach · 2018
Cited alongside, same era.
“Spider: Near-optimal non-convex optimization via stochastic path integrated differential estimator”
Cong Fang, Chris Li, Zhouchen Lin and Tong Zhang · 2018
Cited alongside, same era.
“Neural Tangent Kernel: Convergence and Generalization in Neural Networks”
Arthur Jacot, Franck Gabriel and Clément Hongler · 2018
Cited alongside, same era.
“A mean field view of the landscape of two-layer neural networks”
Song Mei, Andrea Montanari and Phan-Minh Nguyen · 2018
Cited alongside, same era.
“Stochastic Nested Variance Reduction for Nonconvex Optimization”
Dongruo Zhou, Pan Xu and Quanquan Gu · 2018
Cited alongside, same era.
“Learning and Generalization in Overparameterized Neural Networks, Going Beyond Two Layers”
Zeyuan Allen-Zhu, Yuanzhi Li and Yingyu Liang · 2019
Cited alongside, same era.
“Sharp Analysis for Nonconvex SGD Escaping from Saddle Points”
Cong Fang, Zhouchen Lin and Tong Zhang · 2019
Cited alongside, same era.
Yuan Cao and Quanquan Gu · 2020
Later among the works it cites.
“A Generalized Neural Tangent Kernel Analysis for Two-layer Neural Networks”
Zixiang Chen, Yuan Cao, Quanquan Gu and Tong Zhang · 2020
Later among the works it cites.
“Agnostic Learning of a Single Neuron with Gradient Descent”
Spencer Frei, Yuan Cao and Quanquan Gu · 2020
Later among the works it cites.
“Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networks”
Ziwei Ji and Matus Telgarsky · 2020
Later among the works it cites.
“Learning Over-Parametrized Two-Layer ReLU Neural Networks beyond NTK”
Yuanzhi Li, Tengyu Ma and Hongyang. Zhang · 2020
Later among the works it cites.
“Learning a Single Neuron with Gradient Methods”
Gilad Yehudai and Ohad Shamir · 2020
Later among the works it cites.
“Provable Generalization of SGD-trained Neural Networks of Any Width in the Presence of Adversarial Label Noise”
Spencer Frei, Yuan Cao and Quanquan Gu · 2021
Closest in time.
“Loss landscapes and optimization in over-parameterized non-linear systems and neural networks”
Chaoyue Liu, Libin Zhu and Mikhail Belkin · 2021
Closest in time.