Fetching the paper…
Reading the bibliography…
Many modern learning tasks involve fitting nonlinear models to data which are trained in an overparameterized regime where the parameters of the model exceed the size of the training dataset.
A topological property of real analytic subsets
S Lojasiewicz · 1963
Earlier work this paper cites.
Gradient methods for minimizing functionals
Boris Teodorovich Polyak · 1963
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
A nonlinear programming algorithm for solving semidefinite programs via low-rank factorization
Samuel Burer and Renato DC Monteiro · 2003
Earlier work this paper cites.
The generic chaining: upper and lower bounds of stochastic processes
Michel Talagrand · 2006
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Alexander Rakhlin, Ohad Shamir, Karthik Sridharan, et al · 2012
Earlier work this paper cites.
Continuous martingales and Brownian motion
Daniel Revuz and Marc Yor · 2013
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
On the uniform convergence of relative frequencies of events to their probabilities
Vladimir N Vapnik and A Ya Chervonenkis · 2015
Earlier work this paper cites.
Taming the wild: A unified analysis of hogwild-style algorithms
Christopher M De Sa, Ce Zhang, Kunle Olukotun, and Christopher Ré · 2015
Earlier work this paper cites.
Phase retrieval via wirtinger flow: Theory and algorithms
Emmanuel J Candes, Xiaodong Li, and Mahdi Soltanolkotabi · 2015
Earlier work this paper cites.
Low-rank solutions of linear matrix equations via procrustes flow
Stephen Tu, Ross Boczar, Max Simchowitz, Mahdi Soltanolkotabi, and Benjamin Recht · 2015
Earlier work this paper cites.
Tensorflow: a system for large-scale machine learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al · 2016
Earlier work this paper cites.
Global optimality of local search for low rank matrix recovery
Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2016
Earlier work this paper cites.
The non-convex burer-monteiro approach works on smooth semidefinite programs
Nicolas Boumal, Vlad Voroninski, and Afonso Bandeira · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Earlier work this paper cites.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Daniel Soudry and Yair Carmon · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2016
Earlier work this paper cites.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Earlier work this paper cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2017
Earlier work this paper cites.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Earlier work this paper cites.
Geometry of optimization and implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, Ruslan Salakhutdinov, and Nathan Srebro · 2017
Cited alongside, same era.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Cited alongside, same era.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
Alon Brutzkus, Amir Globerson, Eran Malach, and Shai Shalev-Shwartz · 2017
Cited alongside, same era.
Empirical analysis of the hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V Ugur Guney, Yann Dauphin, and Leon Bottou · 2017
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Peter Bartlett, Dylan J. Foster, and Matus Telgarsky · 2017
Spurious valleys in two-layer neural network optimization landscapes
L. Venturi, A. Bandeira, and J. Bruna · 2018
Closest in time.
The global optimization geometry of shallow linear neural networks
Zhihui Zhu, Daniel Soudry, Yonina C. Eldar, and Michael B. Wakin · 2018
Closest in time.
Over-parameterization improves generalization in the xor detection problem
Alon Brutzkus and Amir Globerson · 2018
Closest in time.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Closest in time.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Size-independent sample complexity of neural networks
Noah Golowich, Alexander Rakhlin, and Ohad Shamir · 2017
Cited alongside, same era.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
Alon Brutzkus, Amir Globerson, Eran Malach, and Shai Shalev-Shwartz · 2017
Cited alongside, same era.
Phase retrieval via randomized kaczmarz: theoretical guarantees
Yan Shuo Tan and Roman Vershynin · 2017
Cited alongside, same era.
Structured signal recovery from quadratic measurements: Breaking sample complexity barriers via nonconvex optimization
Mahdi Soltanolkotabi · 2017
Cited alongside, same era.
Siyuan Ma, Raef Bassily, and Mikhail Belkin · 2017
Cited alongside, same era.
Non-convex finite-sum optimization via scsg methods
Lihua Lei, Cheng Ju, Jianbo Chen, and Michael I Jordan · 2017
Cited alongside, same era.
Recovery guarantees for one-hidden-layer neural networks
Kai Zhong, Zhao Song, Prateek Jain, Peter L Bartlett, and Inderjit S Dhillon · 2017
Cited alongside, same era.
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2018
Closest in time.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Closest in time.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2018
Closest in time.
Gradient descent finds global minima of deep neural networks
Simon S Du, Jason D Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2018
Closest in time.
Stronger generalization bounds for deep nets via a compression approach
Sanjeev Arora, Rong Ge, Behnam Neyshabur, and Yi Zhang · 2018
Closest in time.
A mean field view of the landscape of two-layers neural networks
Mei Song, A Montanari, and P Nguyen · 2018
Closest in time.
Learning compact neural networks with regularization
Samet Oymak · 2018
Closest in time.
Overfitting or perfect fitting? risk bounds for classification and regression rules that interpolate
Mikhail Belkin, Daniel Hsu, and Partha Mitra · 2018
Closest in time.
Just interpolate: Kernel "ridgeless" regression can generalize
Tengyuan Liang and Alexander Rakhlin · 2018
Closest in time.
Does data interpolation contradict statistical optimality?
Mikhail Belkin, Alexander Rakhlin, and Alexandre B. Tsybakov · 2018
Closest in time.
Rapid, robust, and reliable blind deconvolution via nonconvex optimization
Xiaodong Li, Shuyang Ling, Thomas Strohmer, and Ke Wei · 2018
Closest in time.
On exponential convergence of sgd in non-convex over-parametrized learning
Raef Bassily, Mikhail Belkin, and Siyuan Ma · 2018
Closest in time.
Fast and faster convergence of sgd for over-parameterized models and an accelerated perceptron
Sharan Vaswani, Francis Bach, and Mark Schmidt · 2018
Closest in time.
A geometric analysis of phase retrieval
Ju Sun, Qing Qu, and John Wright · 2018
Closest in time.
Gradient descent with random initialization: Fast global convergence for nonconvex phase retrieval
Y. Chen, Y. Chi, J. Fan, and C. Ma · 2018
Closest in time.
A theory on the absence of spurious solutions for nonconvex and nonsmooth optimization
Cedric Josz, Yi Ouyang, Richard Y. Zhang, Javad Lavaei, and Somayeh Sojoudi · 2018
Closest in time.
Stochastic gradient descent learns state equations with nonlinear activations
Samet Oymak · 2018
Closest in time.