Fetching the paper…
Reading the bibliography…
We develop a solvable model of neural scaling laws beyond the kernel limit.
Dynamic theory of the spin-glass phase
Haim Sompolinsky and Annette Zippelius · 1981
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Fast rates for regularized least-squares algorithm
Andrea Caponnetto and Ernesto De Vito · 2005
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou · 2017
Earlier work this paper cites.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
Simon S Du, Wei Hu, and Jason D Lee · 2018
Earlier work this paper cites.
Statistical optimality of stochastic gradient descent on hard learning problems through multiple passes
Loucas Pillaud-Vivien, Alessandro Rudi, and Francis Bach · 2018
Earlier work this paper cites.
On lazy training in differentiable programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Earlier work this paper cites.
Similarity of neural network representations revisited
Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey Hinton · 2019
Earlier work this paper cites.
Passed & spurious: Descent algorithms and local minima in spiked matrix-tensor models
Stefano Sarao Mannelli, Florent Krzakala, Pierfrancesco Urbani, and Lenka Zdeborova · 2019
Earlier work this paper cites.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Earlier work this paper cites.
Spectrum dependent learning curves in kernel regression and wide neural networks
Blake Bordelon, Abdulkadir Canatar, and Cengiz Pehlevan · 2020
Earlier work this paper cites.
Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the neural tangent kernel
Stanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani, Daniel M. Roy, and Surya Ganguli · 2020
Earlier work this paper cites.
Disentangling feature and lazy training in deep neural networks
Mario Geiger, Stefano Spigler, Arthur Jacot, and Matthieu Wyart · 2020
Earlier work this paper cites.
When do neural networks outperform kernel methods?
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2020
Earlier work this paper cites.
Statistical field theory for neural networks , volume 970
Moritz Helias and David Dahmen · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Dynamical mean-field theory for stochastic gradient descent in gaussian mixture classification
Francesca Mignacco, Florent Krzakala, Pierfrancesco Urbani, and Lenka Zdeborová · 2020
Earlier work this paper cites.
Towards nngp-guided neural architecture search
Daniel S Park, Jaehoon Lee, Daiyi Peng, Yuan Cao, and Jascha Sohl-Dickstein · 2020
Earlier work this paper cites.
Asymptotic learning curves of kernel methods: empirical data versus teacher–student paradigm
Stefano Spigler, Mario Geiger, and Matthieu Wyart · 2020
Earlier work this paper cites.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Earlier work this paper cites.
The staircase property: How hierarchical structure can guide deep learning
Emmanuel Abbe, Enric Boix-Adsera, Matthew S Brennan, Guy Bresler, and Dheeraj Nagaraj · 2021
Cited alongside, same era.
Explaining neural scaling laws
Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma · 2021
Cited alongside, same era.
Implicit regularization via neural feature alignment
Aristide Baratin, Thomas George, César Laurent, R Devon Hjelm, Guillaume Lajoie, Pascal Vincent, and Simon Lacoste-Julien · 2021
Cited alongside, same era.
Generalization error rates in kernel regression: The crossover from the noiseless to noisy regime
Hugo Cui, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborová · 2021
Cited alongside, same era.
The deep bootstrap framework: Good online learners are good offline generalizers
Preetum Nakkiran, Behnam Neyshabur, and Hanie Sedghi · 2021
Cited alongside, same era.
Sgd learning on neural networks: leap complexity and saddle-to-saddle dynamics
Emmanuel Abbe, Enric Boix Adsera, and Theodor Misiakiewicz · 2023
Later among the works it cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Later among the works it cites.
The onset of variance-limited behavior for networks in the lazy and rich regimes
Alexander Atanasov, Blake Bordelon, Sabarish Sainathan, and Cengiz Pehlevan · 2023
Later among the works it cites.
Depthwise hyperparameter transfer in residual networks: Dynamics and scaling limit, 2023
Blake Bordelon, Lorenzo Noci, Mufan Bill Li, Boris Hanin, and Cengiz Pehlevan · 2023
Later among the works it cites.
Error scaling laws for kernel classification under source and capacity conditions
Hugo Cui, Bruno Loureiro, Florent Krzakala, and Lenka Zdeborová · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Geometric compression of invariant manifolds in neural networks
Jonas Paccolat, Leonardo Petrini, Mario Geiger, Kevin Tyloo, and Matthieu Wyart · 2021
Cited alongside, same era.
Sgd in the large: Average-case analysis, asymptotics, and stepsize criticality
Courtney Paquette, Kiwon Lee, Fabian Pedregosa, and Elliot Paquette · 2021
Cited alongside, same era.
James B Simon, Madeline Dickens, Dhruva Karkada, and Michael R DeWeese · 2021
Cited alongside, same era.
Tensor programs iv: Feature learning in infinite-width neural networks
Greg Yang and Edward J Hu · 2021
Cited alongside, same era.
Tuning large neural networks via zero-shot hyperparameter transfer
Greg Yang, Edward J Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao · 2021
Cited alongside, same era.
Neural networks as kernel learners: The silent alignment effect
Alexander Atanasov, Blake Bordelon, and Cengiz Pehlevan · 2022
Cited alongside, same era.
High-dimensional asymptotics of feature learning: How one gradient step improves the representation
Jimmy Ba, Murat A Erdogdu, Taiji Suzuki, Zhichao Wang, Denny Wu, and Greg Yang · 2022
Cited alongside, same era.
How two-layer neural networks learn, one (giant) step at a time
Yatin Dandi, Florent Krzakala, Bruno Loureiro, Luca Pesce, and Ludovic Stephan · 2023
Later among the works it cites.
Scaling data-constrained language models
Niklas Muennighoff, Alexander M Rush, Boaz Barak, Teven Le Scao, Aleksandra Piktus, Nouamane Tazi, Sampo Pyysalo, Thomas Wolf, and Colin Raffel · 2023
Later among the works it cites.
Neural networks trained with sgd learn distributions of increasing complexity
Maria Refinetti, Alessandro Ingrosso, and Sebastian Goldt · 2023
Later among the works it cites.
James B Simon, Dhruva Karkada, Nikhil Ghosh, and Mikhail Belkin · 2023
Later among the works it cites.
Learning curves for deep structured gaussian feature models, 2023
Jacob A. Zavatone-Veth and Cengiz Pehlevan · 2023
Later among the works it cites.
Scaling and renormalization in high-dimensional regression
Alexander B Atanasov, Jacob A Zavatone-Veth, and Cengiz Pehlevan · 2024
Closest in time.
Learning theory from first principles
Francis Bach · 2024
Closest in time.
Sliding down the stairs: how correlated latent variables accelerate learning with neural networks
Lorenzo Bardone and Sebastian Goldt · 2024
Closest in time.
Infinite-width limit of deep linear neural networks
Lénaïc Chizat, Maria Colombo, Xavier Fernández-Real, and Alessio Figalli · 2024
Closest in time.
Scaling exponents across parameterizations and optimizers
Katie Everett, Lechao Xiao, Mitchell Wortsman, Alexander A Alemi, Roman Novak, Peter J Liu, Izzeddin Gur, Jascha Sohl-Dickstein, Leslie Pack Kaelbling, Jaehoon Lee, et al · 2024
Closest in time.
Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learning
Daniel Kunin, Allan Raventós, Clémentine Dominé, Feng Chen, David Klindt, Andrew Saxe, and Surya Ganguli · 2024
Closest in time.
Scaling laws in linear regression: Compute, parameters, and data
Licong Lin, Jingfeng Wu, Sham M Kakade, Peter L Bartlett, and Jason D Lee · 2024
Closest in time.
4+ 3 phases of compute-optimal neural scaling laws
Elliot Paquette, Courtney Paquette, Lechao Xiao, and Jeffrey Pennington · 2024
Closest in time.
Conditional diffusion mnist
Tim Pearce · 2024
Closest in time.
Mixed dynamics in linear networks: Unifying the lazy and active regimes
Zhenfeng Tu, Santiago Aranguri, and Arthur Jacot · 2024
Closest in time.
Deconstructing what makes a good optimizer for language models
Rosie Zhao, Depen Morwani, David Brandfonbrener, Nikhil Vyas, and Sham Kakade · 2024
Closest in time.