Fetching the paper…
Reading the bibliography…
Large over-parametrized models learned via stochastic gradient descent (SGD) methods have become a key element in modern machine learning.
Gradient methods for minimizing functionals
Boris Teodorovich Polyak · 1963
Earlier work this paper cites.
Gradient methods for convex minimization: better rates under weaker conditions
Hui Zhang and Wotao Yin · 2013
Earlier work this paper cites.
Simple, efficient, and neural algorithms for sparse coding
Sanjeev Arora, Rong Ge, Tengyu Ma, and Ankur Moitra · 2015
Earlier work this paper cites.
Solving random quadratic systems of equations is nearly as easy as solving linear systems
Yuxin Chen and Emmanuel Candes · 2015
Earlier work this paper cites.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Earlier work this paper cites.
Guaranteed matrix completion via non-convex factorization
Ruoyu Sun and Zhi-Quan Luo · 2016
Cited alongside, same era.
Natasha 2: Faster non-convex optimization than sgd
Zeyuan Allen-Zhu · 2017
Cited alongside, same era.
Convergence analysis of two-layer neural networks with relu activation
Yuanzhi Li and Yang Yuan · 2017
Cited alongside, same era.
Diving into the shallows: a computational perspective on large-scale shallow learning
Siyuan Ma and Mikhail Belkin · 2017
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2017
Cited alongside, same era.
Melanie Weber and Suvrit Sra · 2017
Later among the works it cites.
The loss landscape of overparameterized neural networks
Yaim Cooper · 2018
Closest in time.
An alternative view: When does sgd escape local minima?
Robert Kleinberg, Yuanzhi Li, and Yang Yuan · 2018
Closest in time.
The power of interpolation: Understanding the effectiveness of sgd in modern over-parametrized learning
Siyuan Ma, Raef Bassily, and Mikhail Belkin · 2018
Closest in time.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D Lee · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Closest in time.