Fetching the paper…
Reading the bibliography…
In classical statistics, the bias-variance trade-off describes how varying a model's complexity (e.g., number of fit parameters) affects its ability to make accurate predictions.
1903
Earlier work this paper cites.
1906
Earlier work this paper cites.
1912
Earlier work this paper cites.
Vladimir Alexandrovich Marčenko and Leonid Andreevich Pastur, “Distribution of eigenvalues for some sets of random matrices,” Mathematics of the USSR-Sbornik 1
1967
Earlier work this paper cites.
Stuart Geman, Elie Bienenstock, and René Doursat, “Neural Networks and the Bias/Variance Dilemma,” Neural Computation 4
1992
Earlier work this paper cites.
Christopher M. Bishop, Pattern Recognition and Machine Learning (Springer, 2006)
2006
Earlier work this paper cites.
2008
Earlier work this paper cites.
2008
Earlier work this paper cites.
2010
Earlier work this paper cites.
2012
Earlier work this paper cites.
2014
Earlier work this paper cites.
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals, “Understanding Deep Learning Requires Re-thinking Generalization,” International Conference on Learning Representations (ICLR) (2017)
2017
Earlier work this paper cites.
Pankaj Mehta, Marin Bukov, Ching Hao Wang, Alexandre G.R. Day, Clint Richardson, Charles K. Fisher, and David J. Schwab, “A high-bias, low-variance introduction to Machine Learning for physicists,” Physics Reports 810
2019
Earlier work this paper cites.
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal, “Reconciling modern machine-learning practice and the classical bias–variance trade-off,” Proceedings of the National Academy of Sciences 116
2019
Cited alongside, same era.
Jean Barbier, Florent Krzakala, Nicolas Macris, Léo Miolane, and Lenka Zdeborová, “Optimal errors and phase transitions in high-dimensional generalized linear models,” Proceedings of the National Academy of Sciences 116
2019
Cited alongside, same era.
Koby Bibas, Yaniv Fogel, and Meir Feder, “A New Look at an Old Problem: A Universal Learning Approach to Linear Regression,” IEEE International Symposium on Information Theory (ISIT) , 2304–2308 (2019)
2019
Cited alongside, same era.
Andrew K. Lampinen and Surya Ganguli, “An analytic theory of generalization dynamics and transfer learning in deep linear networks,” International Conference on Learning Representations (ICLR) (2019)
2019
Cited alongside, same era.
Zeyu Deng, Abla Kammoun, and Christos Thrampoulidis, “A Model of Double Descent for High-Dimensional Logistic Regression,” IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 4267–4271 (2020)
2020
Later among the works it cites.
Michał Dereziński, Feynman Liang, and Michael W. Mahoney, “Exact expressions for double descent and implicit regularization via surrogate random design,” Advances in Neural Information Processing Systems (NeurIPS) 33
2020
Later among the works it cites.
Federica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mézard, and Lenka Zdeborová, “Generalisation error in learning with random features and the hidden manifold model,” Proceedings of the 37th International Conference on Machine Learning (ICML) PMLR, 119
2020
Later among the works it cites.
Arthur Jacot, Berfin Şimşek, Francesco Spadaro, Clément Hongler, and Franck Gabriel, “Implicit regularization of random feature models,” Proceedings of the 37th International Conference on Machine Learning (ICML) PMLR, 119
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vidya Muthukumar, Kailas Vodrahalli, and Anant Sahai, “Harmless interpolation of noisy data in regression,” IEEE International Symposium on Information Theory (ISIT) , 2299–2303 (2019)
2019
Cited alongside, same era.
Ji Xu and Daniel Hsu, “On the number of variables to use in principal component regression,” Advances in Neural Information Processing Systems (NeurIPS) 32
2019
Cited alongside, same era.
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina, “Entropy-SGD: biasing gradient descent into wide valleys,” Journal of Statistical Mechanics: Theory and Experiment 2019
2019
Cited alongside, same era.
Marco Loog, Tom Viering, Alexander Mey, Jesse H. Krijthe, and David M. J. Tax, “A brief prehistory of double descent,” Proceedings of the National Academy of Sciences 117
2020
Cited alongside, same era.
Ben Adlam and Jeffrey Pennington, “Understanding double descent requires a fine-grained bias-variance decomposition,” Advances in Neural Information Processing Systems (NeurIPS) 33
2020
Cited alongside, same era.
Madhu S. Advani, Andrew M. Saxe, and Haim Sompolinsky, “High-dimensional dynamics of generalization error in neural networks,” Neural Networks 132
2020
Cited alongside, same era.
Jimmy Ba, Murat Erdogdu, Taiji Suzuki, Denny Wu, and Tianzong Zhang, “Generalization of Two-layer Neural Networks: An Asymptotic Viewpoint,” International Conference on Learning Representations (ICLR) (2020)
2020
Cited alongside, same era.
Peter L. Bartlett, Philip M. Long, Gábor Lugosi, and Alexander Tsigler, “Benign overfitting in linear regression,” Proceedings of the National Academy of Sciences 117
2020
Cited alongside, same era.
2020
Later among the works it cites.
Ganesh Ramachandra Kini and Christos Thrampoulidis, “Analytic Study of Double Descent in Binary Classification: The Impact of Loss,” IEEE International Symposium on Information Theory (ISIT) , 2527–2532 (2020)
2020
Later among the works it cites.
Tengyuan Liang and Alexander Rakhlin, “Just interpolate: Kernel “Ridgeless” regression can generalize,” Annals of Statistics 48
2020
Later among the works it cites.
Tengyuan Liang, Alexander Rakhlin, and Xiyu Zhai, “On the Multiple Descent of Minimum-Norm Interpolants and Restricted Lower Isometry of Kernels,” Proceedings of Thirty Third Conference on Learning Theory PMLR, 125
2020
Later among the works it cites.
Zhenyu Liao, Romain Couillet, and Michael W Mahoney, “A random matrix analysis of random Fourier features: beyond the Gaussian kernel, a precise phase transition, and the corresponding double descent,” Advances in Neural Information Processing Systems (NeurIPS) 33
2020
Later among the works it cites.
Zitong Yang, Yaodong Yu, Chong You, Jacob Steinhardt, and Yi Ma, “Rethinking Bias-Variance Trade-off for Generalization of Neural Networks,” Proceedings of the 37th International Conference on Machine Learning (ICML) PMLR, 119
2020
Later among the works it cites.
Carlo Baldassi, Enrico M. Malatesta, Matteo Negri, and Riccardo Zecchina, “Wide flat minima and optimal generalization in classifying high-dimensional Gaussian mixtures,” Journal of Statistical Mechanics: Theory and Experiment 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
Licong Lin and Edgar Dobriban, “What causes the test error? Going beyond bias-variance via ANOVA,” Journal of Machine Learning Research 22
2021
Later among the works it cites.
Song Mei and Andrea Montanari, “The Generalization Error of Random Features Regression: Precise Asymptotics and the Double Descent Curve,” Communications on Pure and Applied Mathematics (2021), 10.1002/cpa.22008
2021
Later among the works it cites.