Fetching the paper…
Reading the bibliography…
Recently, there has been significant progress in understanding the convergence and generalization properties of gradient-based methods for training overparameterized learning models.
The rotation of eigenvectors by a perturbation. III
Chandler Davis and W. M. Kahan · 1970
Earlier work this paper cites.
Cubic regularization of Newton method and its global performance
Yurii Nesterov and Boris T. Polyak · 2006
Earlier work this paper cites.
Trust-region methods
Jorge Nocedal and Stephen J. Wright · 2006
Earlier work this paper cites.
Exact matrix completion via convex optimization
Emmanuel J. Candès and Benjamin Recht · 2009
Earlier work this paper cites.
Smallest singular value of a random rectangular matrix
Mark Rudelson and Roman Vershynin · 2009
Earlier work this paper cites.
Guaranteed minimum-rank solutions of linear matrix equations via nuclear norm minimization
Ben Recht, Maryam Fazel, and Pablo A. Parrilo · 2010
Earlier work this paper cites.
The power of convex relaxation: near-optimal matrix completion
Emmanuel J. Candès and Terence Tao · 2010
Earlier work this paper cites.
Tight oracle inequalities for low-rank matrix recovery from a minimal number of noisy random measurements
Emmanuel J. Candès and Yaniv Plan · 2011
Earlier work this paper cites.
Phaselift: exact and stable signal recovery from magnitude measurements via convex programming
Emmanuel J. Candès, Thomas Strohmer, and Vladislav Voroninski · 2013
Earlier work this paper cites.
Phase retrieval via matrix completion
Emmanuel J. Candès, Yonina C. Eldar, Thomas Strohmer, and Vladislav Voroninski · 2013
Earlier work this paper cites.
Blind deconvolution using convex programming
Ali Ahmed, Benjamin Recht, and Justin Romberg · 2014
Earlier work this paper cites.
Phase retrieval via Wirtinger flow: theory and algorithms
Emmanuel J. Candès, Xiaodong Li, and Mahdi Soltanolkotabi · 2015
Earlier work this paper cites.
Escaping from saddle points — online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Earlier work this paper cites.
An overview of low-rank matrix recovery from incomplete observations
Mark A Davenport and Justin Romberg · 2016
Earlier work this paper cites.
Low-rank solutions of linear matrix equations via procrustes flow
Stephen Tu, Ross Boczar, Max Simchowitz, Mahdi Soltanolkotabi, and Ben Recht · 2016
Earlier work this paper cites.
Matrix completion has no spurious local minimum
Rong Ge, Jason D Lee, and Tengyu Ma · 2016
Earlier work this paper cites.
Gradient descent can take exponential time to escape saddle points
Simon S Du, Chi Jin, Jason D Lee, Michael I Jordan, Aarti Singh, and Barnabas Poczos · 2017
Earlier work this paper cites.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nathan Srebro · 2017
Earlier work this paper cites.
Blind deconvolution meets blind demixing: algorithms and performance bounds
Shuyang Ling and Thomas Strohmer · 2017
Earlier work this paper cites.
Solving random quadratic systems of equations is nearly as easy as solving linear systems
Yuxin Chen and Emmanuel J. Candès · 2017
Earlier work this paper cites.
How to escape saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M. Kakade, and Michael I. Jordan · 2017
Earlier work this paper cites.
Non-convex learning via stochastic gradient langevin dynamics: a nonasymptotic analysis
Maxim Raginsky, Alexander Rakhlin, and Matus Telgarsky · 2017
Earlier work this paper cites.
A hitting time analysis of stochastic gradient langevin dynamics
Yuchen Zhang, Percy Liang, and Moses Charikar · 2017
Earlier work this paper cites.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason D Lee, Daniel Soudry, and Nati Srebro · 2018
Earlier work this paper cites.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2018
Earlier work this paper cites.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
Simon S Du, Wei Hu, and Jason D Lee · 2018
Cited alongside, same era.
High-dimensional probability: An introduction with applications in data science
Roman Vershynin · 2018
Cited alongside, same era.
Gradient descent with identity initialization efficiently learns positive definite linear transformations by deep residual networks
Peter Bartlett, Dave Helmbold, and Philip Long · 2018
Cited alongside, same era.
On the optimization of deep networks: Implicit acceleration by overparameterization
Sanjeev Arora, Nadav Cohen, and Elad Hazan · 2018
Cited alongside, same era.
A convergence analysis of gradient descent for deep linear neural networks
Sanjeev Arora, Nadav Cohen, Noah Golowich, and Wei Hu · 2018
Cited alongside, same era.
Towards resolving the implicit bias of gradient descent for matrix factorization: Greedy low-rank learning
Zhiyuan Li, Yuping Luo, and Kaifeng Lyu · 2021
Later among the works it cites.
Sign-rip: A robust restricted isometry property for low-rank matrix recovery
Jianhao Ma and Salar Fattahi · 2021
Later among the works it cites.
Rank overspecified robust matrix recovery: Subgradient method and exact recovery
Lijun Ding, Liwei Jiang, Yudong Chen, Qing Qu, and Zhihui Zhu · 2021
Later among the works it cites.
More is less: Inducing sparsity via overparameterization
Hung-Hsu Chou, Johannes Maly, and Holger Rauhut · 2021
Later among the works it cites.
Implicit regularization in matrix sensing via mirror descent
Fan Wu and Patrick Rebeschini · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A geometric analysis of phase retrieval
Ju Sun, Qing Qu, and John Wright · 2018
Cited alongside, same era.
On lazy training in differentiable programming
Lénaïc Chizat, Edouard Oyallon, and Francis Bach · 2019
Cited alongside, same era.
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 2019
Cited alongside, same era.
Implicit regularization for optimal sparse recovery
Tomas Vaskevicius, Varun Kanade, and Patrick Rebeschini · 2019
Cited alongside, same era.
Nonconvex optimization meets low-rank matrix factorization: an overview
Yuejie Chi, Yue M. Lu, and Yuxin Chen · 2019
Cited alongside, same era.
Rapid, robust, and reliable blind deconvolution via nonconvex optimization
Xiaodong Li, Shuyang Ling, Thomas Strohmer, and Ke Wei · 2019
Cited alongside, same era.
Regularized gradient descent: a non-convex recipe for fast joint blind deconvolution and demixing
Shuyang Ling and Thomas Strohmer · 2019
Cited alongside, same era.
Later among the works it cites.
Convergence of gradient descent for learning linear neural networks
Gabin Maxime Nguegnang, Holger Rauhut, and Ulrich Terstiege · 2021
Later among the works it cites.
Understanding the dynamics of gradient flow in overparameterized linear models
Salma Tarmoun, Guilherme Franca, Benjamin D Haeffele, and Rene Vidal · 2021
Later among the works it cites.
On the explicit role of initialization on the convergence and implicit bias of overparametrized linear networks
Hancheng Min, Salma Tarmoun, Rene Vidal, and Enrique Mallada · 2021
Later among the works it cites.
Large learning rate tames homogeneity: Convergence and balancing effect
Yuqing Wang, Minshuo Chen, Tuo Zhao, and Molei Tao · 2021
Later among the works it cites.
Beyond procrustes: balancing-free gradient descent for asymmetric low-rank matrix sensing
Cong Ma, Yuanxin Li, and Yuejie Chi · 2021
Later among the works it cites.
On the global convergence of gradient descent for multi-layer resnets in the mean-field regime
Zhiyan Ding, Shi Chen, Qin Li, and Stephen Wright · 2021
Later among the works it cites.
Neural networks can learn representations with gradient descent
Alexandru Damian, Jason Lee, and Mahdi Soltanolkotabi · 2022
Later among the works it cites.
Neural networks efficiently learn low-dimensional representations with sgd
Alireza Mousavi-Hosseini, Sejun Park, Manuela Girotti, Ioannis Mitliagkas, and Murat A Erdogdu · 2022
Later among the works it cites.
Learning single-index models with shallow neural networks
Alberto Bietti, Joan Bruna, Clayton Sanford, and Min Jae Song · 2022
Later among the works it cites.
Algorithmic regularization in model-free overparametrized asymmetric matrix factorization
Liwei Jiang, Yudong Chen, and Lijun Ding · 2022
Later among the works it cites.
A validation approach to over-parameterized matrix and image recovery
Lijun Ding, Zhen Qin, Liwei Jiang, Jinxin Zhou, and Zhihui Zhu · 2022
Later among the works it cites.
Flat minima generalize for low-rank matrix recovery
Lijun Ding, Dmitriy Drusvyatskiy, and Maryam Fazel · 2022
Later among the works it cites.
Non-negative least squares via overparametrization
Hung-Hsu Chou, Johannes Maly, and Claudio Mayrink Verdun · 2022
Later among the works it cites.
Implicit bias of gradient descent on reparametrized models: On equivalence to mirror descent
Zhiyuan Li, Tianhao Wang, JasonD Lee, and Sanjeev Arora · 2022
Later among the works it cites.
Learning deep linear neural networks: Riemannian gradient flows and convergence to global minimizers
Bubacarr Bah, Holger Rauhut, Ulrich Terstiege, and Michael Westdickenberg · 2022
Later among the works it cites.
Proof methods for robust low-rank matrix recovery
Tim Fuchs, David Gross, Peter Jung, Felix Krahmer, Richard Kueng, and Dominik Stöger · 2022
Later among the works it cites.
Randomly initialized alternating least squares: Fast convergence for matrix sensing
Kiryung Lee and Dominik Stöger · 2022
Later among the works it cites.
Understanding incremental learning of gradient descent: A fine-grained analysis of matrix sensing
Jikai Jin, Zhiyuan Li, Kaifeng Lyu, Simon S Du, and Jason D Lee · 2023
Closest in time.
The power of preconditioning in overparameterized low-rank matrix sensing
Xingyu Xu, Yandi Shen, Yuejie Chi, and Cong Ma · 2023
Closest in time.
A line-search descent algorithm for strict saddle functions with complexity guarantees
Michael O’Neill and Stephen J Wright · 2023
Closest in time.