Fetching the paper…
Reading the bibliography…
For small training set sizes $P$, the generalization error of wide neural networks is well-approximated by the error of an infinite width neural network (NN), either in the kernel or mean-field/feature-learning regime.
Algorithms for learning kernels based on centered alignment
Corinna Cortes, Mehryar Mohri, and Afshin Rostamizadeh · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Wide residual networks, 2017
Sergey Zagoruyko and Nikos Komodakis · 2017
Earlier work this paper cites.
Jax: composable transformations of python+ numpy programs
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, et al · 2018
Earlier work this paper cites.
Neural tangent kernel: convergence and generalization in neural networks (invited paper)
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Earlier work this paper cites.
On lazy training in differentiable programming
Lénaïc Chizat, Edouard Oyallon, and Francis R. Bach · 2019
Earlier work this paper cites.
Finite depth and width corrections to the neural tangent kernel
Boris Hanin and Mihai Nica · 2019
Earlier work this paper cites.
Similarity of neural network representations revisited
Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey Hinton · 2019
Earlier work this paper cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel S. Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jascha Sohl-Dickstein · 2019
Earlier work this paper cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc Le · 2019
Earlier work this paper cites.
Spectrum dependent learning curves in kernel regression and wide neural networks
Blake Bordelon, Abdulkadir Canatar, and Cengiz Pehlevan · 2020
Earlier work this paper cites.
Triple descent and the two kinds of overfitting: Where & why do they appear?
Stéphane d’Ascoli, Levent Sagun, and Giulio Biroli · 2020
Earlier work this paper cites.
A precise performance analysis of learning with random features
Oussama Dhifallah and Yue M Lu · 2020
Earlier work this paper cites.
Asymptotics of wide networks from feynman diagrams
Ethan Dyer and Guy Gur-Ari · 2020
Earlier work this paper cites.
Double trouble in double descent: Bias and variance (s) in the lazy regime
Stéphane d’Ascoli, Maria Refinetti, Giulio Biroli, and Florent Krzakala · 2020
Cited alongside, same era.
Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the neural tangent kernel
Stanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani, Daniel M Roy, and Surya Ganguli · 2020
Cited alongside, same era.
Scaling description of generalization with number of parameters in deep learning
Mario Geiger, Arthur Jacot, Stefano Spigler, Franck Gabriel, Levent Sagun, Stéphane d’Ascoli, Giulio Biroli, Clément Hongler, and Matthieu Wyart · 2020
Cited alongside, same era.
Disentangling feature and lazy training in deep neural networks
Mario Geiger, Stefano Spigler, Arthur Jacot, and Matthieu Wyart · 2020
Cited alongside, same era.
Generalisation error in learning with random features and the hidden manifold model
Federica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mézard, and Lenka Zdeborová · 2020
Cited alongside, same era.
Bruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt, Florent Krzakala, Marc Mézard, and Lenka Zdeborová · 2021
Later among the works it cites.
What can linearized neural networks actually say about generalization?
Guillermo Ortiz-Jiménez, Seyed-Mohsen Moosavi-Dezfooli, and Pascal Frossard · 2021
Later among the works it cites.
Geometric compression of invariant manifolds in neural networks
Jonas Paccolat, Leonardo Petrini, Mario Geiger, Kevin Tyloo, and Matthieu Wyart · 2021
Later among the works it cites.
Geometric compression of invariant manifolds in neural networks
Jonas Paccolat, Leonardo Petrini, Mario Geiger, Kevin Tyloo, and Matthieu Wyart · 2021
Later among the works it cites.
The principles of deep learning theory, 2021
Daniel A. Roberts, Sho Yaida, and Boris Hanin · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
When do neural networks outperform kernel methods?
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2020
Cited alongside, same era.
Universality laws for high-dimensional learning with random features
Hong Hu and Yue M Lu · 2020
Cited alongside, same era.
Neural tangents: Fast and easy infinite neural networks in python
Roman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee, Alexander A. Alemi, Jascha Sohl-Dickstein, and Samuel S. Schoenholz · 2020
Cited alongside, same era.
Asymptotic learning curves of kernel methods: empirical data versus teacher–student paradigm
Stefano Spigler, Mario Geiger, and Matthieu Wyart · 2020
Cited alongside, same era.
Feature learning in infinite-width neural networks
Greg Yang and Edward J. Hu · 2020
Cited alongside, same era.
Neural networks as kernel learners: The silent alignment effect
Alexander Atanasov, Blake Bordelon, and Cengiz Pehlevan · 2021
Cited alongside, same era.
Explaining neural scaling laws, 2021
Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma · 2021
Cited alongside, same era.
Neural tangent kernel eigenvalues accurately predict generalization, 2021
James B. Simon, Madeline Dickens, and Michael R. DeWeese · 2021
Later among the works it cites.
High-dimensional asymptotics of feature learning: How one gradient step improves the representation
Jimmy Ba, Murat A Erdogdu, Taiji Suzuki, Zhichao Wang, Denny Wu, and Greg Yang · 2022
Closest in time.
Self-consistent dynamical field theory of kernel evolution in wide neural networks
Blake Bordelon and Cengiz Pehlevan · 2022
Closest in time.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Closest in time.
A solvable model of neural scaling laws, 2022
Alexander Maloney, Daniel A. Roberts, and James Sully · 2022
Closest in time.
The generalization error of random features regression: Precise asymptotics and the double descent curve
Song Mei and Andrea Montanari · 2022
Closest in time.
Limitations of the ntk for understanding generalization in deep learning
Nikhil Vyas, Yamini Bansal, and Preetum Nakkiran · 2022
Closest in time.
More than a toy: Random matrix models predict how real-world neural representations generalize
Alexander Wei, Wei Hu, and Jacob Steinhardt · 2022
Closest in time.
Contrasting random and learned features in deep bayesian linear regression
Jacob A Zavatone-Veth, William L Tong, and Cengiz Pehlevan · 2022
Closest in time.