Fetching the paper…
Reading the bibliography…
Deep networks are typically trained with many more parameters than the size of the training dataset.
Two models of double descent for weak features
Belkin, M.; Hsu, D.; and Xu, J. 2019 · 1903
Earlier work this paper cites.
Surprises in high-dimensional ridgeless least squares interpolation
Hastie, T.; Montanari, A.; Rosset, S.; and Tibshirani, R. J. 2019 · 1903
Earlier work this paper cites.
The generalization error of random features regression: Precise asymptotics and double descent curve
Mei, S.; and Montanari, A. 2019 · 1908
Earlier work this paper cites.
A Model of Double Descent for High-dimensional Binary Linear Classification
Deng, Z.; Kammoun, A.; and Thrampoulidis, C. 2019 · 1911
Earlier work this paper cites.
Montanari, A.; Ruan, F.; Sohn, Y.; and Yan, J. 2019 · 1911
Earlier work this paper cites.
Exact expressions for double descent and implicit regularization via surrogate random design
Dereziński, M.; Liang, F.; and Mahoney, M. W. 2019 · 1912
Earlier work this paper cites.
Deep double descent: Where bigger models and more data hurt
Nakkiran, P.; Kaplun, G.; Bansal, Y.; Yang, T.; Barak, B.; and Sutskever, I. 2019 · 1912
Earlier work this paper cites.
Minimax theorems
Fan, K. 1953 · 1953
Earlier work this paper cites.
Cox’s regression model for counting processes: a large sample study
Andersen, P. K.; and Gill, R. D. 1982 · 1982
Earlier work this paper cites.
Optimal brain damage
LeCun, Y.; Denker, J. S.; and Solla, S. A. 1990 · 1990
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Hassibi, B.; and Stork, D. G. 1993 · 1993
Earlier work this paper cites.
Optimal brain surgeon: Extensions and performance comparisons
Hassibi, B.; Stork, D. G.; and Wolff, G. 1994 · 1994
Earlier work this paper cites.
Large sample estimation and hypothesis testing
Newey, W. K.; and McFadden, D. 1994 · 1994
Earlier work this paper cites.
Analytic study of double descent in binary classification: The impact of loss
Kini, G.; and Thrampoulidis, C. 2020 · 2001
Earlier work this paper cites.
A precise high-dimensional asymptotic theory for boosting and min-l1-norm interpolated classifiers
Liang, T.; and Sur, P. 2020 · 2002
Earlier work this paper cites.
Proving the Lottery Ticket Hypothesis: Pruning is All You Need
Malach, E.; Yehudai, G.; Shalev-Shwartz, S.; and Shamir, O. 2020 · 2002
Earlier work this paper cites.
Picking winning tickets before training by preserving gradient flow
Wang, C.; Zhang, G.; and Grosse, R. 2020 · 2002
Earlier work this paper cites.
The Gaussian equivalence of generative models for learning with two-layer neural networks
Goldt, S.; Reeves, G.; Mézard, M.; Krzakala, F.; and Zdeborová, L. 2020 · 2006
Earlier work this paper cites.
Exploring Weight Importance and Hessian Bias in Model Pruning
Li, M.; Sattar, Y.; Thrampoulidis, C.; and Oymak, S. 2020 · 2006
Earlier work this paper cites.
Optimal Lottery Tickets via SubsetSum: Logarithmic Over-Parameterization is Sufficient
Pensia, A.; Rajput, S.; Nagle, A.; Vishwakarma, H.; and Papailiopoulos, D. 2020 · 2006
Earlier work this paper cites.
Fundamental limits of ridge-regularized empirical risk minimization in high dimensions
Taheri, H.; Pedarsani, R.; and Thrampoulidis, C. 2020 · 2006
Earlier work this paper cites.
Universality laws for high-dimensional learning with random features
Hu, H.; and Lu, Y. M. 2020 · 2009
Cited alongside, same era.
Benign overfitting in ridge regression
Tsigler, A.; and Bartlett, P. L. 2020 · 2009
Cited alongside, same era.
The dynamics of message passing on dense graphs, with applications to compressed sensing
Bayati, M.; and Montanari, A. 2011 · 2011
Cited alongside, same era.
State evolution for general approximate message passing algorithms, with applications to spatial coupling
Javanmard, A.; and Montanari, A. 2013 · 2013
Cited alongside, same era.
The squared-error of generalized lasso: A precise analysis
Oymak, S.; Thrampoulidis, C.; and Hassibi, B. 2013 · 2013
Cited alongside, same era.
Snip: Single-shot network pruning based on connection sensitivity
Lee, N.; Ajanthan, T.; and Torr, P. H. 2018 · 2018
Later among the works it cites.
Just interpolate: Kernel" ridgeless" regression can generalize
Liang, T.; and Rakhlin, A. 2018 · 2018
Later among the works it cites.
The distribution of the lasso: Uniform control over sparse balls and adaptive parameter tuning
Miolane, L.; and Montanari, A. 2018 · 2018
Later among the works it cites.
Learning Compact Neural Networks with Regularization
Oymak, S. 2018 · 2018
Later among the works it cites.
Universality laws for randomized dimension reduction, with applications
Oymak, S.; and Tropp, J. A. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stojnic, M. 2013 · 2013
Cited alongside, same era.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Neyshabur, B.; Tomioka, R.; and Srebro, N. 2014 · 2014
Cited alongside, same era.
Han, S.; Mao, H.; and Dally, W. J. 2015 · 2015
Cited alongside, same era.
Learning both weights and connections for efficient neural network
Han, S.; Pool, J.; Tran, J.; and Dally, W. 2015 · 2015
Cited alongside, same era.
Regularized linear regression: A precise analysis of the estimation error
Thrampoulidis, C.; Oymak, S.; and Hassibi, B. 2015 · 2015
Cited alongside, same era.
Training skinny deep neural networks with iterative hard thresholding methods
Jin, X.; Yuan, X.; Feng, J.; and Yan, S. 2016 · 2016
Cited alongside, same era.
Implicit regularization in matrix factorization
Gunasekar, S.; Woodworth, B. E.; Bhojanapalli, S.; Neyshabur, B.; and Srebro, N. 2017 · 2017
Cited alongside, same era.
Mobilenetv2: Inverted residuals and linear bottlenecks
Sandler, M.; Howard, A.; Zhu, M.; Zhmoginov, A.; and Chen, L.-C. 2018 · 2018
Later among the works it cites.
Precise Error Analysis of Regularized M M -Estimators in High Dimensions
Thrampoulidis, C.; Abbasi, E.; and Hassibi, B. 2018 · 2018
Later among the works it cites.
Symbol error rate performance of box-relaxation decoders in massive mimo
Thrampoulidis, C.; Xu, W.; and Hassibi, B. 2018 · 2018
Later among the works it cites.
Universality in learning from linear measurements
Abbasi, E.; Salehi, F.; and Hassibi, B. 2019 · 2019
Later among the works it cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Belkin, M.; Hsu, D.; Ma, S.; and Mandal, S. 2019 · 2019
Later among the works it cites.
Does data interpolation contradict statistical optimality?
Belkin, M.; Rakhlin, A.; and Tsybakov, A. B. 2019 · 2019
Later among the works it cites.
On lazy training in differentiable programming
Chizat, L.; Oyallon, E.; and Bach, F. 2019 · 2019
Later among the works it cites.
The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
Frankle, J.; and Carbin, M. 2019 · 2019
Later among the works it cites.
The impact of regularization on high-dimensional logistic regression
Salehi, F.; Abbasi, E.; and Hassibi, B. 2019 · 2019
Later among the works it cites.
Mnasnet: Platform-aware neural architecture search for mobile
Tan, M.; Chen, B.; Pang, R.; Vasudevan, V.; Sandler, M.; Howard, A.; and Le, Q. V. 2019 · 2019
Later among the works it cites.
Deconstructing lottery tickets: Zeros, signs, and the supermask
Zhou, H.; Lan, J.; Liu, R.; and Yosinski, J. 2019 · 2019
Later among the works it cites.
Benign overfitting in linear regression
Bartlett, P. L.; Long, P. M.; Lugosi, G.; and Tsigler, A. 2020 · 2020
Closest in time.
Overfitting Can Be Harmless for Basis Pursuit, But Only to a Degree
Ju, P.; Lin, X.; and Liu, J. 2020 · 2020
Closest in time.
Towards moderate overparameterization: global convergence guarantees for training shallow neural networks
Oymak, S.; and Soltanolkotabi, M. 2020 · 2020
Closest in time.
The Performance Analysis of Generalized Margin Maximizers on Separable Data
Salehi, F.; Abbasi, E.; and Hassibi, B. 2020 · 2020
Closest in time.