Fetching the paper…
Reading the bibliography…
We show that the gradient descent algorithm provides an implicit regularization effect in the learning of over-parameterized matrix factorization models and one-hidden-layer neural networks with quadratic activations.
Gradient methods for minimizing functionals
Boris Teodorovich Polyak · 1963
Earlier work this paper cites.
A simple weight decay can improve generalization
Anders Krogh and John A Hertz · 1992
Earlier work this paper cites.
Stability and generalization
Olivier Bousquet and André Elisseeff · 2002
Earlier work this paper cites.
Rank, trace-norm and max-norm
Nathan Srebro and Adi Shraibman · 2005
Earlier work this paper cites.
Learning theory: stability is sufficient for generalization and necessary and sufficient for consistency of empirical risk minimization
Sayan Mukherjee, Partha Niyogi, Tomaso Poggio, and Ryan Rifkin · 2006
Earlier work this paper cites.
The restricted isometry property and its implications for compressed sensing
Emmanuel J Candes · 2008
Earlier work this paper cites.
Exact matrix completion via convex optimization
Emmanuel J Candès and Benjamin Recht · 2009
Earlier work this paper cites.
Quantitative estimates of the convergence of the empirical covariance matrix in log-concave ensembles
Radosław Adamczak, Alexander Litvak, Alain Pajor, and Nicole Tomczak-Jaegermann · 2010
Earlier work this paper cites.
Guaranteed minimum-rank solutions of linear matrix equations via nuclear norm minimization
Benjamin Recht, Maryam Fazel, and Pablo A Parrilo · 2010
Earlier work this paper cites.
Learnability, stability and uniform convergence
Shai Shalev-Shwartz, Ohad Shamir, Nathan Srebro, and Karthik Sridharan · 2010
Earlier work this paper cites.
Robust principal component analysis?
Emmanuel J Candès, Xiaodong Li, Yi Ma, and John Wright · 2011
Earlier work this paper cites.
A simpler approach to matrix completion
Benjamin Recht · 2011
Earlier work this paper cites.
On the universality of online mirror descent
Nati Srebro, Karthik Sridharan, and Ambuj Tewari · 2011
Earlier work this paper cites.
Weighted low-rank approximations
Nathan Srebro and Tommi Jaakkola · 2013
Earlier work this paper cites.
On the computational efficiency of training neural networks
Roi Livni, Shai Shalev-Shwartz, and Ohad Shamir · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Cited alongside, same era.
Exact and stable covariance estimation from quadratic sampling via convex programming
Yuxin Chen, Yuejie Chi, and Andrea J Goldsmith · 2015
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Benjamin Recht, and Yoram Singer · 2015
Cited alongside, same era.
Low-rank solutions of linear matrix equations via Procrustes flow
Stephen Tu, Ross Boczar, Mahdi Soltanolkotabi, and Benjamin Recht · 2015
Cited alongside, same era.
Parseval networks: Improving robustness to adversarial examples
Moustapha Cisse, Piotr Bojanowski, Edouard Grave, Yann Dauphin, and Nicolas Usunier · 2017
Closest in time.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Closest in time.
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Closest in time.
On the Optimization Landscape of Tensor Decompositions
R. Ge and T. Ma · 2017
Closest in time.
No spurious local minima in nonconvex low rank problems: A unified geometric analysis
Rong Ge, Chi Jin, and Yi Zheng · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kai Zhong, Prateek Jain, and Inderjit S Dhillon · 2015
Cited alongside, same era.
Global Optimality of Local Search for Low Rank Matrix Recovery
S. Bhojanapalli, B. Neyshabur, and N. Srebro · 2016
Cited alongside, same era.
Matrix completion has no spurious local minimum
Rong Ge, Jason D. Lee, and Tengyu Ma · 2016
Cited alongside, same era.
Gradient descent learns linear dynamical systems
Moritz Hardt, Tengyu Ma, and Benjamin Recht · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
No bad local minima: Data independent training error guarantees for multilayer neural networks
Daniel Soudry and Yair Carmon · 2016
Cited alongside, same era.
A geometric analysis of phase retrieval
Ju Sun, Qing Qu, and John Wright · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nathan Srebro · 2017
Closest in time.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2017
Closest in time.
Kronecker recurrent units
Cijo Jose, Moustpaha Cisse, and Francois Fleuret · 2017
Closest in time.
Low rank matrix recovery from rank one measurements
Richard Kueng, Holger Rauhut, and Ulrich Terstiege · 2017
Closest in time.
A pac-bayesian approach to spectrally-normalized margin bounds for neural networks
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nathan Srebro · 2017
Closest in time.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, and Nati Srebro · 2017
Closest in time.
Theoretical insights into the optimization landscape of over-parameterized shallow neural networks
M. Soltanolkotabi, A. Javanmard, and J. D. Lee · 2017
Closest in time.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, and Nathan Srebro · 2017
Closest in time.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nathan Srebro, and Benjamin Recht · 2017
Closest in time.
On the learnability of fully-connected neural networks
Yuchen Zhang, Jason Lee, Martin Wainwright, and Michael Jordan · 2017
Closest in time.