Fetching the paper…
Reading the bibliography…
Modern machine learning classifiers often exhibit vanishing classification error on the training set.
Thomas M Cover, Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition , IEEE Transactions on Electronic Computers (1965), no. 3, 326–334
1965
Earlier work this paper cites.
Marc Mézard, Giorgio Parisi, and Miguel A. Virasoro, Spin glass theory and beyond , World Scientific, 1987
1987
Earlier work this paper cites.
Elizabeth Gardner, The space of interactions in neural network models , Journal of physics A: Mathematical and general 21
1988
Earlier work this paper cites.
Yehoram Gordon, On Milman’s inequality and random subspaces which escape through a mesh in R n R^{n} , Geometric Aspects of Functional Analysis, Springer, 1988, pp. 84–106
1988
Earlier work this paper cites.
Radford M Neal, Priors for infinite networks , Bayesian Learning for Neural Networks, Springer, 1996, pp. 29–53
1996
Earlier work this paper cites.
Peter L Bartlett, The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network , IEEE Transactions on Information Theory 44
1998
Earlier work this paper cites.
Andreas Engel and Christian Van den Broeck, Statistical mechanics of learning , Cambridge University Press, 2001
2001
Earlier work this paper cites.
Peter L Bartlett and Shahar Mendelson, Rademacher and gaussian complexities: Risk bounds and structural results , Journal of Machine Learning Research 3
2002
Earlier work this paper cites.
Vladimir Koltchinskii and Dmitry Panchenko, Empirical margin distributions and bounding the generalization error of combined classifiers , The Annals of Statistics 30
2002
Earlier work this paper cites.
Grace Wahba, Soft and hard classification by reproducing kernel hilbert space methods , Proceedings of the National Academy of Sciences 99
2002
Earlier work this paper cites.
Mariya Shcherbina and Brunello Tirozzi, Rigorous solution of the Gardner problem , Communications in Mathematical Physics 234
2003
Earlier work this paper cites.
Maria-Florina Balcan, Avrim Blum, and Santosh Vempala, Kernels as features: On kernels, margins, and low-dimensional mappings , Machine Learning 65
2006
Earlier work this paper cites.
T Hofmann, B Schölkopf, and AJ Smola, Kernel methods in machine learning , The Annals of Statistics 36
2008
Earlier work this paper cites.
Vladimir Koltchinskii, Oracle inequalities in empirical risk minimization and sparse recovery problems , vol. 2033, Springer Science & Business Media, 2011, Ecole d’Eté de Probabilités de Saint-Flour XXXVIII-2008
2008
Earlier work this paper cites.
Ali Rahimi and Benjamin Recht, Random features for large-scale kernel machines , Advances in Neural Information Processing Systems, 2008, pp. 1177–1184
2008
Earlier work this paper cites.
Cédric Villani, Optimal transport: old and new , vol. 338, Springer Science & Business Media, 2008
2008
Earlier work this paper cites.
Martin Anthony and Peter L Bartlett, Neural network learning: Theoretical foundations , Cambridge University Press, 2009
2009
Earlier work this paper cites.
Sham M Kakade, Karthik Sridharan, and Ambuj Tewari, On the complexity of linear prediction: Risk bounds, margin bounds, and regularization , Advances in neural information processing systems, 2009, pp. 793–800
2009
Earlier work this paper cites.
Marc Mézard and Andrea Montanari, Information, Physics and Computation , Oxford, 2009
2009
Earlier work this paper cites.
Z. Bai and J. Silverstein, Spectral Analysis of Large Dimensional Random Matrices , Springer, 2010
2010
Earlier work this paper cites.
Mihailo Stojnic, l1 optimization and its various thresholds in compressed sensing , 2010 IEEE International Conference on Acoustics, Speech and Signal Processing, 2010, pp. 3910–3913
2010
Earlier work this paper cites.
Mohsen Bayati and Andrea Montanari, The dynamics of message passing on dense graphs, with applications to compressed sensing , IEEE Trans. on Inform. Theory 57
2011
Earlier work this paper cites.
Peter Bühlmann and Sara Van De Geer, Statistics for high-dimensional data: methods, theory and applications , Springer Science & Business Media, 2011
2011
Earlier work this paper cites.
V. Chandrasekaran, B. Recht, P. A. Parrilo, and A.S. Willsky, The convex geometry of linear inverse problems , Foundations of Computational Mathematics 12
2012
Earlier work this paper cites.
Xiuyuan Cheng and Amit Singer, The spectrum of random inner-product kernel matrices , Random Matrices: Theory and Applications 2
2013
Earlier work this paper cites.
David L Donoho, Iain Johnstone, and Andrea Montanari, Accurate prediction of phase transitions in compressed sensing via a connection to minimax denoising , IEEE transactions on information theory 59
2013
Earlier work this paper cites.
Noureddine El Karoui, Derek Bean, Peter J Bickel, Chinghway Lim, and Bin Yu, On robust regression with high-dimensional predictors , Proceedings of the National Academy of Sciences 110
2013
Earlier work this paper cites.
Jean-Baptiste Hiriart-Urruty and Claude Lemaréchal, Convex analysis and minimization algorithms i: Fundamentals , vol. 305, Springer science & business media, 2013
2013
Earlier work this paper cites.
Samet Oymak, Christos Thrampoulidis, and Babak Hassibi, The squared-error of generalized lasso: A precise analysis , 2013 51st Annual Allerton Conference on Communication, Control, and Computing (Allerton), 2013, pp. 1002–1009
2013
Earlier work this paper cites.
Dennis Amelunxen, Martin Lotz, Michael B McCoy, and Joel A Tropp, Living on the edge: Phase transitions in convex programs with random data , Information and Inference: A Journal of the IMA 3
2014
Cited alongside, same era.
Shai Shalev-Shwartz and Shai Ben-David, Understanding machine learning: From theory to algorithms , Cambridge University Press, 2014
2014
Cited alongside, same era.
Mohsen Bayati, Marc Lelarge, and Andrea Montanari, Universality in polytope phase transitions and message passing algorithms , The Annals of Applied Probability 25
2015
Cited alongside, same era.
Christos Thrampoulidis, Samet Oymak, and Babak Hassibi, Regularized linear regression: A precise analysis of the estimation error , Conference on Learning Theory, 2015, pp. 1683–1709
2015
Cited alongside, same era.
Jean Barbier and Nicolas Macris, The adaptive interpolation method: a simple scheme to prove replica formulas in bayesian inference , Probability Theory and Related Fields 174
2019
Closest in time.
Zhou Fan and Andrea Montanari, The spectral norm of random inner-product kernel matrices , Probability Theory and Related Fields 173
2019
Closest in time.
2019
Closest in time.
Marc Lelarge and Léo Miolane, Fundamental limits of symmetric low-rank matrix estimation , Probability Theory and Related Fields 173
2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
David Donoho and Andrea Montanari, High dimensional robust m-estimation: Asymptotic variance via approximate message passing , Probability Theory and Related Fields 166
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2017
Cited alongside, same era.
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nati Srebro, Exploring generalization in deep learning , Advances in neural information processing systems 30
2017
Cited alongside, same era.
2018
Cited alongside, same era.
Mikhail Belkin, Daniel J Hsu, and Partha Mitra, Overfitting or perfect fitting? risk bounds for classification and regression rules that interpolate , Advances in Neural Information Processing Systems, 2018, pp. 2300–2311
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.
Pragya Sur and Emmanuel J Candès, A modern maximum-likelihood theory for high-dimensional logistic regression , Proceedings of the National Academy of Sciences 116
2019
Closest in time.
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler, Benign overfitting in linear regression , Proceedings of the National Academy of Sciences 117
2020
Closest in time.
Federica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mézard, and Lenka Zdeborová, Generalisation error in learning with random features and the hidden manifold model , International Conference on Machine Learning, PMLR, 2020, pp. 3452–3462
2020
Closest in time.
Ganesh Ramachandra Kini and Christos Thrampoulidis, Analytic study of double descent in binary classification: The impact of loss , 2020 IEEE International Symposium on Information Theory (ISIT), IEEE, 2020, pp. 2527–2532
2020
Closest in time.
Tengyuan Liang and Alexander Rakhlin, Just interpolate: Kernel “ridgeless” regression can generalize , The Annals of Statistics 48
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
Niladri S Chatterji and Philip M Long, Finite-sample analysis of interpolating linear classifiers in the overparameterized regime , The Journal of Machine Learning Research 22
2021
Closest in time.
Ganesh Ramachandra Kini, Orestis Paraskevas, Samet Oymak, and Christos Thrampoulidis, Label-imbalanced and group-sensitive classification under overparameterization , Advances in Neural Information Processing Systems 34
2021
Closest in time.
Frederic Koehler, Lijia Zhou, Danica J Sutherland, and Nathan Srebro, Uniform convergence of interpolators: Gaussian width, norm bounds and benign overfitting , Advances in Neural Information Processing Systems 34
2021
Closest in time.
Vidya Muthukumar, Adhyyan Narang, Vignesh Subramanian, Mikhail Belkin, Daniel Hsu, and Anant Sahai, Classification vs regression in overparameterized regimes: Does the loss function matter? , Journal of Machine Learning Research 22
2021
Closest in time.
Zeyu Deng, Abla Kammoun, and Christos Thrampoulidis, A model of double descent for high-dimensional binary linear classification , Information and Inference: A Journal of the IMA 11
2022
Closest in time.
Spencer Frei, Niladri S Chatterji, and Peter Bartlett, Benign overfitting without linearity: Neural network classifiers trained by gradient descent for noisy linear data , Conference on Learning Theory, PMLR, 2022, pp. 2668–2703
2022
Closest in time.
Hong Hu and Yue M Lu, Universality laws for high-dimensional learning with random features , IEEE Transactions on Information Theory (2022)
2022
Closest in time.
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani, Surprises in high-dimensional ridgeless least squares interpolation , The Annals of Statistics 50
2022
Closest in time.
Adel Javanmard and Mahdi Soltanolkotabi, Precise statistical analysis of classification accuracies for adversarial training , The Annals of Statistics 50
2022
Closest in time.
Tengyuan Liang and Pragya Sur, A precise high-dimensional asymptotic theory for boosting and minimum- ℓ 1 \ell_{1} -norm interpolated classifiers , The Annals of Statistics 50
2022
Closest in time.
Andrea Montanari and Basil N Saeed, Universality of empirical risk minimization , Conference on Learning Theory, PMLR, 2022, pp. 4310–4312
2022
Closest in time.
Andrea Montanari and Yiqiao Zhong, The interpolation phase transition in neural networks: Memorization and generalization under lazy training , The Annals of Statistics 50
2022
Closest in time.
2022
Closest in time.
2023
Closest in time.
Andrea Montanari, , Feng Ruan, Basil N Saeed, and Youngtak Sohn, Universality of max-margin classification , 2023, In preparation
2023
Closest in time.