Fetching the paper…
Reading the bibliography…
This work characterizes the effect of depth on the optimization landscape of linear regression, showing that, despite their nonconvexity, deeper models have more desirable optimization landscape.
Weak convergence and empirical processes: with applications to statistics
Aad W Van Der Vaart, Adrianus Willem van der Vaart, Aad van der Vaart, and Jon Wellner · 1996
Earlier work this paper cites.
The elements of statistical learning
Jerome Friedman, Trevor Hastie, Robert Tibshirani, et al · 2001
Earlier work this paper cites.
A nonlinear programming algorithm for solving semidefinite programs via low-rank factorization
Samuel Burer and Renato DC Monteiro · 2003
Earlier work this paper cites.
The dantzig selector: Statistical estimation when p is much larger than n
Emmanuel Candes and Terence Tao · 2007
Earlier work this paper cites.
The deterministic lasso
Sara Van de Geer · 2007
Earlier work this paper cites.
Supremum concentration inequality and modulus of continuity for sub-nth chaos processes
Frederi G Viens and Andrew B Vizcarra · 2007
Earlier work this paper cites.
A simple proof of the restricted isometry property for random matrices
Richard Baraniuk, Mark Davenport, Ronald DeVore, and Michael Wakin · 2008
Earlier work this paper cites.
Simultaneous analysis of lasso and dantzig selector
Peter J Bickel, Ya’acov Ritov, and Alexandre B Tsybakov · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Minimax rates of estimation for high-dimensional linear regression over ℓ q \ell_{q} -balls
Garvesh Raskutti, Martin J Wainwright, and Bin Yu · 2011
Earlier work this paper cites.
Representation benefits of deep feedforward networks
Matus Telgarsky · 2015
Earlier work this paper cites.
Best subset selection via a modern optimization lens
Dimitris Bertsimas, Angela King, and Rahul Mazumder · 2016
Earlier work this paper cites.
Global optimality of local search for low rank matrix recovery
Srinadh Bhojanapalli, Behnam Neyshabur, and Nathan Srebro · 2016
Earlier work this paper cites.
The power of depth for feedforward neural networks
Ronen Eldan and Ohad Shamir · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Cited alongside, same era.
Benefits of depth in neural networks
Matus Telgarsky · 2016
Cited alongside, same era.
Subset selection with shrinkage: Sparse linear modeling when the snr is low
Rahul Mazumder, Peter Radchenko, and Antoine Dedieu · 2017
Cited alongside, same era.
Damek Davis and Dmitriy Drusvyatskiy · 2018
Cited alongside, same era.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
Simon S Du, Wei Hu, and Jason D Lee · 2018
Benign overfitting in linear regression
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler · 2020
Later among the works it cites.
Exact guarantees on the absence of spurious local minima for non-negative rank-1 robust principal component analysis
Salar Fattahi and Somayeh Sojoudi · 2020
Later among the works it cites.
Evaluation of neural architectures trained with square loss vs cross-entropy in classification tasks
Like Hui and Mikhail Belkin · 2020
Later among the works it cites.
Nonconvex robust low-rank matrix recovery
Xiao Li, Zhihui Zhu, Anthony Man-Cho So, and Rene Vidal · 2020
Later among the works it cites.
Zhiyuan Li, Yuping Luo, and Kaifeng Lyu · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro · 2018
Cited alongside, same era.
Sensitivity and generalization in neural networks: an empirical study
Roman Novak, Yasaman Bahri, Daniel A Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Cited alongside, same era.
Implicit regularization in deep matrix factorization
Sanjeev Arora, Nadav Cohen, Wei Hu, and Yuping Luo · 2019
Cited alongside, same era.
Width provably matters in optimization for deep linear neural networks
Simon Du and Wei Hu · 2019
Cited alongside, same era.
The implicit bias of depth: How incremental learning drives generalization
Daniel Gissin, Shai Shalev-Shwartz, and Amit Daniely · 2019
Cited alongside, same era.
Implicit regularization for optimal sparse recovery
Tomas Vaskevicius, Varun Kanade, and Patrick Rebeschini · 2019
Cited alongside, same era.
High-dimensional statistics: A non-asymptotic viewpoint
Martin J Wainwright · 2019
Cited alongside, same era.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Later among the works it cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Later among the works it cites.
More is less: Inducing sparsity via overparameterization
Hung-Hsu Chou, Johannes Maly, and Holger Rauhut · 2021
Later among the works it cites.
Rank overspecified robust matrix recovery: Subgradient method and exact recovery
Lijun Ding, Liwei Jiang, Yudong Chen, Qing Qu, and Zhihui Zhu · 2021
Later among the works it cites.
Shape matters: Understanding the implicit bias of the noise covariance
Jeff Z HaoChen, Colin Wei, Jason Lee, and Tengyu Ma · 2021
Later among the works it cites.
Implicit sparse regularization: The impact of depth and early stopping
Jiangyuan Li, Thanh Nguyen, Chinmay Hegde, and Ka Wai Wong · 2021
Later among the works it cites.
Sign-rip: A robust restricted isometry property for low-rank matrix recovery
Jianhao Ma and Salar Fattahi · 2021
Later among the works it cites.
Optimization-based separations for neural networks
Itay Safran and Jason D Lee · 2021
Later among the works it cites.
Jianhao Ma and Salar Fattahi · 2022
Closest in time.