Fetching the paper…
Reading the bibliography…
The population loss of trained deep neural networks often follows precise power-law scaling relations with either the size of the training dataset or the number of parameters in the network.
Das asymptotische verteilungsgesetz der eigenwerte linearer partieller differentialgleichungen (mit einer anwendung auf die theorie der hohlraumstrahlung)
Hermann Weyl · 1912
Earlier work this paper cites.
A sequence of approximated solutions to the SK model for spin glasses
Giorgio Parisi · 1980
Earlier work this paper cites.
Eigenvalues of positive definite kernels
JB Reade · 1983
Earlier work this paper cites.
Eigenvalues of integral operators with smooth positive definite kernels
Thomas Kühn · 1987
Earlier work this paper cites.
Scaling and generalization in neural networks: a case study
Subutai Ahmad and Gerald Tesauro · 1989
Earlier work this paper cites.
Generalized Linear Models , volume 37
P McCullagh and John A Nelder · 1989
Earlier work this paper cites.
Can neural networks do better than the vapnik-chervonenkis bounds?
David Cohn and Gerald Tesauro · 1991
Earlier work this paper cites.
Bayesian Learning for Neural Networks
Radford M. Neal · 1994
Earlier work this paper cites.
Learning curves for gaussian processes
Peter Sollich · 1998
Earlier work this paper cites.
Interpolation of Spatial Data: Some Theory for Kriging
Michael L Stein · 1999
Earlier work this paper cites.
Upper and lower bounds on the learning curve for gaussian processes
Christopher KI Williams and Francesco Vivarelli · 2000
Earlier work this paper cites.
A variational approach to learning curves
Dörthe Malzahn and Manfred Opper · 2001
Earlier work this paper cites.
Learning curves for gaussian processes regression: A framework for good approximations
Dörthe Malzahn and Manfred Opper · 2001
Earlier work this paper cites.
Learning curves for gaussian process regression: Approximations and bounds
Peter Sollich and Anason Halees · 2002
Earlier work this paper cites.
Learning curves and bootstrap estimates for inference with gaussian processes: A statistical mechanics study
Dörthe Malzahn and Manfred Opper · 2003
Earlier work this paper cites.
Maximum likelihood estimation of intrinsic dimension
Elizaveta Levina and Peter J Bickel · 2005
Earlier work this paper cites.
Optimal rates for the regularized least-squares algorithm
Andrea Caponnetto and Ernesto De Vito · 2007
Earlier work this paper cites.
Local polynomial regression on unknown manifolds
Peter J Bickel, Bo Li, et al · 2007
Earlier work this paper cites.
Notes on regularized least squares, 2007
Ryan M Rifkin and Ross A Lippert · 2007
Earlier work this paper cites.
Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning
Ali Rahimi and Benjamin Recht · 2008
Earlier work this paper cites.
Optimal rates for regularized least squares regression
Ingo Steinwart, Don R Hush, Clint Scovel, et al · 2009
Earlier work this paper cites.
Eigenvalues of integral operators defined by smooth positive definite kernels
JC Ferreira and VA Menegatto · 2009
Earlier work this paper cites.
The elements of statistical learning: data mining, inference, and prediction
Trevor Hastie, Robert Tibshirani, and Jerome Friedman · 2009
Earlier work this paper cites.
Approximating manifolds by meshes: asymptotic bounds in higher codimension
David de Laat · 2011
Earlier work this paper cites.
Replica theory for learning curves for gaussian processes on random graphs
Matthew J Urry and Peter Sollich · 2012
Earlier work this paper cites.
Wide residual networks
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md Patwary, Mostofa Ali, Yang Yang, and Yanqi Zhou · 2017
Earlier work this paper cites.
High-dimensional dynamics of generalization error in neural networks
Madhu S Advani and Andrew M Saxe · 2017
Cited alongside, same era.
How close are the eigenvectors of the sample and actual covariance matrices?
Andreas Loukas · 2017
Cited alongside, same era.
Deep information propagation
Samuel S Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein · 2017
Cited alongside, same era.
Deep neural networks as Gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Sam Schoenholz, Jeffrey Pennington, and Jascha Sohl-dickstein · 2018
Cited alongside, same era.
Gaussian process behaviour in wide deep neural networks
Alexander G. de G. Matthews, Jiri Hron, Mark Rowland, Richard E. Turner, and Zoubin Ghahramani · 2018
Cited alongside, same era.
Neural Tangent Kernel: Convergence and generalization in neural networks
Sobolev norm learning rates for regularized least-squares algorithms
Simon Fischer and Ingo Steinwart · 2020
Later among the works it cites.
Double trouble in double descent: Bias and variance (s) in the lazy regime
Stéphane d’Ascoli, Maria Refinetti, Giulio Biroli, and Florent Krzakala · 2020
Later among the works it cites.
High-dimensional dynamics of generalization error in neural networks
Madhu S Advani, Andrew M Saxe, and Haim Sompolinsky · 2020
Later among the works it cites.
Asymptotics of wide convolutional neural networks
Anders Andreassen and Ethan Dyer · 2020
Later among the works it cites.
Scaling description of generalization with number of parameters in deep learning
Mario Geiger, Arthur Jacot, Stefano Spigler, Franck Gabriel, Levent Sagun, Stéphane d’Ascoli, Giulio Biroli, Clément Hongler, and Matthieu Wyart · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Cited alongside, same era.
Dynamical isometry and a mean field theory of CNNs: How to train 10,000-layer vanilla convolutional neural networks
Lechao Xiao, Yasaman Bahri, Jascha Sohl-Dickstein, Samuel Schoenholz, and Jeffrey Pennington · 2018
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel S. Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Cited alongside, same era.
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani · 2019
Cited alongside, same era.
The generalization error of random features regression: Precise asymptotics and double descent curve
Song Mei and Andrea Montanari · 2019
Cited alongside, same era.
More data can hurt for linear regression: Sample-wise double descent
Preetum Nakkiran · 2019
Cited alongside, same era.
Non-Gaussian processes and neural networks at finite widths
Sho Yaida · 2020
Later among the works it cites.
Finite depth and width corrections to the neural tangent kernel
Boris Hanin and Mihai Nica · 2020
Later among the works it cites.
The large learning rate phase of deep learning: the catapult mechanism
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2020
Later among the works it cites.
Implicit bias of deep linear networks in the large learning rate phase
Wei Huang, Weitao Du, Richard Yi Da Xu, and Chunrui Liu · 2020
Later among the works it cites.
Distributional generalization: A new kind of generalization
Preetum Nakkiran and Yamini Bansal · 2020
Later among the works it cites.
Your classifier is secretly an energy based model and you should treat it like one
Will Grathwohl, Kuan-Chieh Wang, Joern-Henrik Jacobsen, David Duvenaud, Mohammad Norouzi, and Kevin Swersky · 2020
Later among the works it cites.
Neural Tangents: Fast and easy infinite neural networks in python
Roman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee, Alexander A. Alemi, Jascha Sohl-Dickstein, and Samuel S. Schoenholz · 2020
Later among the works it cites.
Neural kernels without tangents
Vaishaal Shankar, Alex Chengyu Fang, Wenshuo Guo, Sara Fridovich-Keil, Ludwig Schmidt, Jonathan Ragan-Kelley, and Benjamin Recht · 2020
Later among the works it cites.
Flax: A neural network library and ecosystem for JAX, 2020
Jonathan Heek, Anselm Levskaya, Avital Oliver, Marvin Ritter, Bertrand Rondepierre, Andreas Steiner, and Marc van Zee · 2020
Later among the works it cites.
Generalisation error in learning with random features and the hidden manifold model
Federica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mézard, and Lenka Zdeborová · 2020
Later among the works it cites.
Finite versus infinite neural networks: an empirical study
Jaehoon Lee, Samuel Schoenholz, Jeffrey Pennington, Ben Adlam, Lechao Xiao, Roman Novak, and Jascha Sohl-Dickstein · 2020
Later among the works it cites.
On the predictability of pruning across scales
Jonathan S Rosenfeld, Jonathan Frankle, Michael Carbin, and Nir Shavit · 2021
Closest in time.
A theoretical-empirical approach to estimating sample complexity of dnns
Devansh Bisla, Apoorva Nandini Saridena, and Anna Choromanska · 2021
Closest in time.
Marcus Hutter · 2021
Closest in time.
Generalization error rates in kernel regression: The crossover from the noiseless to noisy regime
Hugo Cui, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborová · 2021
Closest in time.
The deep bootstrap framework: Good online learners are good offline generalizers
Preetum Nakkiran, Behnam Neyshabur, and Hanie Sedghi · 2021
Closest in time.
Spectral bias and task-model alignment explain generalization in kernel regression and infinitely wide neural networks
Abdulkadir Canatar, Blake Bordelon, and Cengiz Pehlevan · 2021
Closest in time.
Learning curves for overparametrized deep neural networks: A field theory perspective
Omry Cohen, Or Malka, and Zohar Ringel · 2021
Closest in time.
An empirical analysis of compute-optimal large language model training
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Closest in time.
Scaling laws from the data manifold dimension
Utkarsh Sharma and Jared Kaplan · 2022
Closest in time.
A solvable model of neural scaling laws
Alexander Maloney, Daniel A Roberts, and James Sully · 2022
Closest in time.
More than a toy: Random matrix models predict how real-world neural representations generalize
Alexander Wei, Wei Hu, and Jacob Steinhardt · 2022
Closest in time.