Fetching the paper…
Reading the bibliography…
From benign overfitting in overparameterized models to rich power-law scalings in performance, simple ridge regression displays surprising behaviors sometimes thought to be limited to deep neural networks.
1909
Earlier work this paper cites.
1912
Earlier work this paper cites.
Hu, Hong, and Yue M Lu (2022b), “Universality laws for high-dimensional learning with random features,” IEEE Transactions on Information Theory 69
1964
Earlier work this paper cites.
Widom, Benjamin (1965), “Equation of state in the neighborhood of the critical point,” The Journal of Chemical Physics 43
1965
Earlier work this paper cites.
Kadanoff, Leo P (1966), “Scaling laws for Ising models near T c {T}_{c} ,” Physics Physique Fizika 2
1966
Earlier work this paper cites.
Kadanoff, Leo P, Wolfgang Götze, David Hamblen, Robert Hecht, EAS Lewis, V V_ Palciauskas, Martin Rayl, J Swift, David Aspnes, and Joseph Kane (1967), “Static phenomena near critical points: theory and experiment,” Reviews of Modern Physics 39
1967
Earlier work this paper cites.
Marchenko, Vladimir Alexandrovich, and Leonid Andreevich Pastur (1967), “Distribution of eigenvalues for some sets of random matrices,” Matematicheskii Sbornik 114
1967
Earlier work this paper cites.
Wilson, Kenneth G, and John Kogut (1974), “The renormalization group and the ϵ \epsilon expansion,” Physics reports 12
1974
Earlier work this paper cites.
Craven, Peter, and Grace Wahba (1978), “Smoothing noisy data with spline functions: estimating the correct degree of smoothing by the method of generalized cross-validation,” Numerische mathematik 31
1978
Earlier work this paper cites.
Weingarten, Don (1978), “Asymptotic behavior of group integrals in the limit of infinite rank,” Journal of Mathematical Physics 19
1978
Earlier work this paper cites.
Golub, Gene H, Michael Heath, and Grace Wahba (1979), “Generalized cross-validation as a method for choosing a good ridge parameter,” Technometrics 21
1979
Earlier work this paper cites.
Fahrmeir, Ludwig, and Heinz Kaufmann (1985), “Consistency and Asymptotic Normality of the Maximum Likelihood Estimator in Generalized Linear Models,” The Annals of Statistics 13
1985
Earlier work this paper cites.
Ahmad, Subutai, and Gerald Tesauro (1988), “Scaling and generalization in neural networks: a case study,” Advances in neural information processing systems 1
1988
Earlier work this paper cites.
Neudecker, H, and A.M. Wesselman (1990), “The asymptotic variance matrix of the sample correlation matrix,” Linear Algebra and its Applications 127
1990
Earlier work this paper cites.
Krogh, Anders, and John A Hertz (1992), “Generalization in a linear perceptron in the presence of noise,” Journal of Physics A: Mathematical and General 25
1992
Earlier work this paper cites.
Voiculescu, Dan V, Ken J Dykema, and Alexandru Nica (1992), Free random variables (American Mathematical Society)
1992
Earlier work this paper cites.
Watkin, Timothy L H, Albrecht Rau, and Michael Biehl (1993), “The statistical mechanics of learning a rule,” Rev. Mod. Phys. 65
1993
Earlier work this paper cites.
Brouwer, PW, and CWJ Beenakker (1996), “Diagrammatic method of integration over the unitary group, with applications to quantum transport in mesoscopic systems,” Journal of Mathematical Physics 37
1996
Earlier work this paper cites.
Cardy, John (1996), Scaling and renormalization in statistical physics , Vol. 5 (Cambridge university press)
1996
Earlier work this paper cites.
Ruderman, Daniel L (1997), “Origins of scaling in natural images,” Vision research 37
1997
Earlier work this paper cites.
Voiculescu, Dan V (1997), Free probability theory , Vol. 12 (American Mathematical Soc.)
1997
Earlier work this paper cites.
Sollich, Peter (1998), “Learning curves for Gaussian processes,” Advances in neural information processing systems 11
1998
Earlier work this paper cites.
Cramér, Harald (1999), Mathematical methods of statistics , Vol. 26 (Princeton university press)
1999
Earlier work this paper cites.
Dietrich, Rainer, Manfred Opper, and Haim Sompolinsky (1999), “Statistical mechanics of support vector networks,” Physical review letters 82
1999
Earlier work this paper cites.
Engel, Andreas, and Christian van den Broeck (2001), Statistical Mechanics of Learning (Cambridge University Press)
2001
Earlier work this paper cites.
2001
Earlier work this paper cites.
Schölkopf, Bernhard, and Alexander J Smola (2002), Learning with kernels: support vector machines, regularization, optimization, and beyond (MIT press)
2002
Earlier work this paper cites.
Sollich, Peter, and Anason Halees (2002), “Learning curves for Gaussian process regression: Approximations and bounds,” Neural computation 14
2002
Earlier work this paper cites.
Burda, Zdzisław, Jerzy Jurkiewicz, and Bartłomiej Wacław (2005), “Spectral moments of correlated wishart matrices,” Phys. Rev. E 71
2005
Earlier work this paper cites.
Caponnetto, Andrea, and Ernesto De Vito (2005), Fast rates for regularized least-squares algorithm , Tech. Rep. (Massachusetts Institute of Technology Computer Science and Artificial Intelligence Laboratory)
2005
Earlier work this paper cites.
2006
Earlier work this paper cites.
Nica, Alexandru, and Roland Speicher (2006), Lectures on the combinatorics of free probability , Vol. 13 (Cambridge University Press)
2006
Earlier work this paper cites.
Williams, Christopher KI, and Carl Edward Rasmussen (2006), Gaussian processes for machine learning (MIT press Cambridge, MA)
2006
Earlier work this paper cites.
Caponnetto, Andrea, and Ernesto De Vito (2007), “Optimal rates for the regularized least-squares algorithm,” Foundations of Computational Mathematics 7
2007
Earlier work this paper cites.
2008
Earlier work this paper cites.
Collins, Benoît, and Sho Matsumoto (2009), “On some properties of orthogonal Weingarten functions,” Journal of Mathematical Physics 50
2009
Earlier work this paper cites.
Hastie, Trevor, Robert Tibshirani, Jerome H Friedman, and Jerome H Friedman (2009), The elements of statistical learning: data mining, inference, and prediction , Vol. 2 (Springer)
2009
Earlier work this paper cites.
Hyvärinen, Aapo, Jarmo Hurri, and Patrick O Hoyer (2009), Natural image statistics: A probabilistic approach to early computational vision. , Vol. 39 (Springer Science & Business Media)
2009
Earlier work this paper cites.
Steinwart, Ingo, Don R Hush, Clint Scovel, et al. (2009), “Optimal rates for regularized least squares regression.” in COLT , pp. 79–93
2009
Earlier work this paper cites.
Banica, Teodor (2010), “The orthogonal Weingarten formula in compact form,” Letters in Mathematical Physics 91
2010
Earlier work this paper cites.
Burda, Z, A. Jarosz, G. Livan, M. A. Nowak, and A. Swiech (2010), “Eigenvalues and singular values of products of rectangular Gaussian random matrices,” Physical Review E 82
2010
Earlier work this paper cites.
2010
Earlier work this paper cites.
Burda, Z, RA Janik, and MA Nowak (2011), “Multiplication law and S S transform for non-Hermitian random matrices,” Physical Review E 84
2011
Earlier work this paper cites.
Horn, Roger A, and Charles R Johnson (2012), Matrix Analysis (Cambridge University Press)
2012
Earlier work this paper cites.
Tao, Terence, and Van Vu (2014), “Random matrices: the universality phenomenon for Wigner ensembles,” Modern Aspects of Random Matrix Theory 72
2014
Earlier work this paper cites.
Bun, Joël, Romain Allez, Jean-Philippe Bouchaud, and Marc Potters (2016), “Rotational invariant estimator for general noisy matrices,” IEEE Transactions on Information Theory 62
2016
Earlier work this paper cites.
Dicker, Lee H (2016), “Ridge regression and asymptotic minimax estimation over spheres of growing dimension,” Bernoulli 22
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
Mingo, James A, and Roland Speicher (2017), Free probability and random matrices , Vol. 35 (Springer)
2017
Earlier work this paper cites.
Pennington, Jeffrey, and Pratik Worah (2017), “Nonlinear random matrix theory for deep learning,” Advances in neural information processing systems 30
2017
Cited alongside, same era.
Vaswani, Ashish, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin (2017), “Attention is all you need,” Advances in neural information processing systems 30
2017
Cited alongside, same era.
Belkin, Mikhail, Siyuan Ma, and Soumik Mandal (2018), “To understand deep learning we need to understand kernel learning,” in Proceedings of the 35th International Conference on Machine Learning , Proceedings of Machine Learning Research, Vol. 80, edited by Jennifer Dy and Andreas Krause (PMLR) pp. 541–549
2018
Cited alongside, same era.
Dobriban, Edgar, and Stefan Wager (2018), “High-dimensional asymptotics of prediction: Ridge regression and classification,” The Annals of Statistics 46
2018
Cited alongside, same era.
2022
Later among the works it cites.
Lee, Kiwon, Andrew Cheng, Elliot Paquette, and Courtney Paquette (2022), “Trajectory of mini-batch momentum: Batch size saturation and convergence in high dimensions,” in Advances in Neural Information Processing Systems , Vol. 35, edited by S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Curran Associates, Inc.) pp. 36944–36957
2022
Later among the works it cites.
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Louart, Cosme, Zhenyu Liao, and Romain Couillet (2018), “A random matrix approach to neural networks,” The Annals of Applied Probability 28
2018
Cited alongside, same era.
Mei, Song, Andrea Montanari, and Phan-Minh Nguyen (2018), “A mean field view of the landscape of two-layer neural networks,” Proceedings of the National Academy of Sciences 115
2018
Cited alongside, same era.
Peskin, Michael E (2018), An Introduction to quantum field theory (CRC press)
2018
Cited alongside, same era.
Pillaud-Vivien, Loucas, Alessandro Rudi, and Francis Bach (2018), “Statistical optimality of stochastic gradient descent on hard learning problems through multiple passes,” Advances in Neural Information Processing Systems 31
2018
Cited alongside, same era.
Radford, Alec, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. (2018), “Improving language understanding by generative pre-training,”
2018
Cited alongside, same era.
Ali, Alnur, J. Zico Kolter, and Ryan J. Tibshirani (2019), “A continuous-time view of early stopping for least squares regression,” in Proceedings of the Twenty-Second International Conference on Artificial Intelligence and Statistics , Proceedings of Machine Learning Research, Vol. 89, edited by Kamalika Chaudhuri and Masashi Sugiyama (PMLR) pp. 1370–1378
2019
Cited alongside, same era.
Belkin, Mikhail, Daniel Hsu, Siyuan Ma, and Soumik Mandal (2019), “Reconciling modern machine-learning practice and the classical bias–variance trade-off,” Proceedings of the National Academy of Sciences 116
2019
Cited alongside, same era.
Chizat, Lenaic, Edouard Oyallon, and Francis Bach (2019), “On lazy training in differentiable programming,” Advances in neural information processing systems 32
2019
Cited alongside, same era.
2022
Later among the works it cites.
Mei, Song, Theodor Misiakiewicz, and Andrea Montanari (2022), “Generalization error of random feature and kernel methods: Hypercontractivity and kernel matrix concentration,” Applied and Computational Harmonic Analysis 59
2022
Later among the works it cites.
Mei, Song, and Andrea Montanari (2022), “The generalization error of random features regression: Precise asymptotics and the double descent curve,” Communications on Pure and Applied Mathematics 75
2022
Later among the works it cites.
2022
Later among the works it cites.
Montanari, Andrea, and Basil N. Saeed (2022), “Universality of empirical risk minimization,” in Proceedings of Thirty Fifth Conference on Learning Theory , Proceedings of Machine Learning Research, Vol. 178, edited by Po-Ling Loh and Maxim Raginsky (PMLR) pp. 4310–4312
2022
Later among the works it cites.
Paquette, Courtney, Elliot Paquette, Ben Adlam, and Jeffrey Pennington (2022), “Implicit regularization or implicit conditioning? exact risk trajectories of sgd in high dimensions,” in Advances in Neural Information Processing Systems , Vol. 35, edited by S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Curran Associates, Inc.) pp. 35984–35999
2022
Later among the works it cites.
Roberts, Daniel A, Sho Yaida, and Boris Hanin (2022), The principles of deep learning theory , Vol. 46 (Cambridge University Press Cambridge, MA, USA)
2022
Later among the works it cites.
Rocks, Jason W, and Pankaj Mehta (2022), “Bias-variance decomposition of overparameterized regression with random linear features,” Physical Review E 106
2022
Later among the works it cites.
Sharma, Utkarsh, and Jared Kaplan (2022), “Scaling laws from the data manifold dimension,” Journal of Machine Learning Research 23
2022
Later among the works it cites.
Tomasini, Umberto M, Antonio Sclocchi, and Matthieu Wyart (2022), “Failure and success of the spectral bias prediction for Laplace kernel ridge regression: the case of low-dimensional data,” in International Conference on Machine Learning (PMLR) pp. 21548–21583
2022
Later among the works it cites.
Wei, Alexander, Wei Hu, and Jacob Steinhardt (2022), “More than a toy: Random matrix models predict how real-world neural representations generalize,” in International Conference on Machine Learning (PMLR) pp. 23549–23588
2022
Later among the works it cites.
Xiao, Lechao, Hong Hu, Theodor Misiakiewicz, Yue Lu, and Jeffrey Pennington (2022), “Precise learning curves and higher-order scalings for dot-product kernel regression,” in Advances in Neural Information Processing Systems , Vol. 35, edited by S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Curran Associates, Inc.) pp. 4558–4570
2022
Later among the works it cites.
Zavatone-Veth, Jacob A, Abdulkadir Canatar, Benjamin S Ruben, and Cengiz Pehlevan (2022a), “Asymptotics of representation learning in finite Bayesian neural networks,” Journal of Statistical Mechanics: Theory and Experiment 2022
2022
Later among the works it cites.
Zhai, Xiaohua, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer (2022), “Scaling vision transformers,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pp. 12104–12113
2022
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Bordelon, Blake, and Cengiz Pehlevan (2023), “Dynamics of finite width kernel and prediction fluctuations in mean field neural networks,” in Advances in Neural Information Processing Systems , Vol. 36, edited by A. Oh, T. Neumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Curran Associates, Inc.) pp. 9707–9750
2023
Later among the works it cites.
Cui, Hugo, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborová (2023), “Error scaling laws for kernel classification under source and capacity conditions,” Machine Learning: Science and Technology 4
2023
Later among the works it cites.
Dandi, Yatin, Ludovic Stephan, Florent Krzakala, Bruno Loureiro, and Lenka Zdeborová (2023), “Universality laws for Gaussian mixtures in generalized linear models,” in Advances in Neural Information Processing Systems , Vol. 36, edited by A. Oh, T. Neumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Curran Associates, Inc.) pp. 54754–54768
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Pesce, Luca, Florent Krzakala, Bruno Loureiro, and Ludovic Stephan (2023), “Are Gaussian data all you need? The extents and limits of universality in high-dimensional generalized linear estimation,” in Proceedings of the 40th International Conference on Machine Learning , Proceedings of Machine Learning Research, Vol. 202, edited by Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (PMLR) pp. 27680–27708
2023
Later among the works it cites.
2023
Later among the works it cites.
Simon, James B, Madeline Dickens, Dhruva Karkada, and Michael Deweese (2023), “The eigenlearning framework: A conservation law perspective on kernel ridge regression and wide neural networks,” Transactions on Machine Learning Research
2023
Later among the works it cites.
Tao, Terence (2023), Topics in random matrix theory , Vol. 132 (American Mathematical Society)
2023
Later among the works it cites.
Alabdulmohsin, Ibrahim M, Xiaohua Zhai, Alexander Kolesnikov, and Lucas Beyer (2024), “Getting ViT in shape: Scaling laws for compute-optimal model design,” Advances in Neural Information Processing Systems 36
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Bach, Francis (2024), “High-dimensional analysis of double descent for linear regression with random projections,” SIAM Journal on Mathematics of Data Science 6
2024
Closest in time.
Bachmann, Gregor, Sotiris Anagnostidis, and Thomas Hofmann (2024), “Scaling MLPs: A tale of inductive bias,” Advances in Neural Information Processing Systems 36
2024
Closest in time.
Bahri, Yasaman, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma (2024), “Explaining neural scaling laws,” Proceedings of the National Academy of Sciences 121
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Canatar, Abdulkadir, Jenelle Feather, Albert Wakhloo, and SueYeon Chung (2024), “A spectral theory of neural prediction and alignment,” Advances in Neural Information Processing Systems 36
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Michaud, Eric, Ziming Liu, Uzay Girit, and Max Tegmark (2024), “The quantization model of neural scaling,” Advances in Neural Information Processing Systems 36
2024
Closest in time.
2024
Closest in time.
Muennighoff, Niklas, Alexander Rush, Boaz Barak, Teven Le Scao, Nouamane Tazi, Aleksandra Piktus, Sampo Pyysalo, Thomas Wolf, and Colin A Raffel (2024), “Scaling data-constrained language models,” Advances in Neural Information Processing Systems 36
2024
Closest in time.
Patil, Pratik, and Daniel LeJeune (2024), “Asymptotically free sketched ridge ensembles: Risks, cross-validation, and tuning,” in The Twelfth International Conference on Learning Representations
2024
Closest in time.
2024
Closest in time.
Vyas, Nikhil, Alexander Atanasov, Blake Bordelon, Depen Morwani, Sabarish Sainathan, and Cengiz Pehlevan (2024), “Feature-learning networks are consistent across widths at realistic scales,” Advances in Neural Information Processing Systems 36
2024
Closest in time.
Muller, Ralf R (2002), “On the asymptotic eigenvalue distribution of concatenated vector-valued fading channels,” IEEE Transactions on Information Theory 48
2091
Closest in time.