Fetching the paper…
Reading the bibliography…
Generalization beyond a training dataset is a main goal of machine learning, but theoretical understanding of generalization remains an open problem for many models.
Spin glass theory and beyond: An Introduction to the Replica Method and Its Applications
Marc Mézard, Giorgio Parisi, and Miguel Virasoro · 1987
Earlier work this paper cites.
The space of interactions in neural network models
Elizabeth Gardner · 1988
Earlier work this paper cites.
Phase transitions in simple learning
J A Hertz, A Krogh, and G I Thorbergsson · 1989
Earlier work this paper cites.
Spline models for observational data
Grace Wahba · 1990
Earlier work this paper cites.
On the ability of the optimal perceptron to generalise
M Opper, W Kinzel, J Kleinz, and R Nehl · 1990
Earlier work this paper cites.
Statistical mechanics of learning from examples
H. S. Seung, H. Sompolinsky, and N. Tishby · 1992
Earlier work this paper cites.
The statistical mechanics of learning a rule
Timothy L. H. Watkin, Albrecht Rau, and Michael Biehl · 1993
Earlier work this paper cites.
Finite-size effects in learning and generalization in linear perceptrons
Peter Sollich · 1994
Earlier work this paper cites.
Bayesian learning for neural networks
Radford M Neal · 1996
Earlier work this paper cites.
Statistical Learning Theory
Vladimir N. Vapnik · 1998
Earlier work this paper cites.
Generalization in a linear perceptron in the presence of noise
A Krogh and J. Hertz · 1999
Earlier work this paper cites.
Learning curves for gaussian processes
Peter Sollich · 1999
Earlier work this paper cites.
Statistical mechanics of support vector networks
Rainer Dietrich, Manfred Opper, and Haim Sompolinsky · 1999
Earlier work this paper cites.
Statistical mechanics of support vector networks
Rainer Dietrich, Manfred Opper, and Haim Sompolinsky · 1999
Earlier work this paper cites.
Regularization networks and support vector machines
Theodoros Evgeniou, Massimiliano Pontil, and Tomaso Poggio · 2000
Earlier work this paper cites.
Learning curves for gaussian processes regression: A framework for good approximations
Dörthe Malzahn and Manfred Opper · 2001
Earlier work this paper cites.
Statistical mechanics of learning
Andreas Engel and Christian Van den Broeck · 2001
Earlier work this paper cites.
Classes of kernels for machine learning: a statistics perspective
Marc G Genton · 2001
Earlier work this paper cites.
Best choices for regularization parameters in learning theory: on the bias-variance problem
Felipe Cucker and Steve Smale · 2002
Earlier work this paper cites.
Learning with kernels: support vector machines, regularization, optimization, and beyond
Bernhard Schölkopf, Alexander J Smola, Francis Bach, et al · 2002
Earlier work this paper cites.
On kernel-target alignment
Nello Cristianini, John Shawe-Taylor, André Elisseeff, and Jaz Kandol a · 2002
Earlier work this paper cites.
Kernel methods for pattern analysis
John Shawe-Taylor, Nello Cristianini, et al · 2004
Earlier work this paper cites.
Gaussian Processes for Machine Learning (Adaptive Computation and Machine Learning)
Carl Edward Rasmussen and Christopher K. I. Williams · 2005
Earlier work this paper cites.
Spin-glass theory for pedestrians
Tommaso Castellani and Andrea Cavagna · 2005
Earlier work this paper cites.
Optimal rates for the regularized least-squares algorithm
Andrea Caponnetto and Ernesto De Vito · 2007
Earlier work this paper cites.
The elements of statistical learning: data mining, inference, and prediction
Trevor Hastie, Robert Tibshirani, and Jerome Friedman · 2009
Earlier work this paper cites.
Kernel methods for deep learning
Youngmin Cho and Lawrence K Saul · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Cited alongside, same era.
Mnist handwritten digit database
Yann LeCun, Corinna Cortes, and CJ Burges · 2010
Cited alongside, same era.
Spectral analysis of large dimensional random matrices
Zhidong Bai and Jack W Silverstein · 2010
Cited alongside, same era.
Algorithms for learning kernels based on centered alignment
Corinna Cortes, Mehryar Mohri, and Afshin Rostamizadeh · 2012
Cited alongside, same era.
Approximation Theory and Harmonic Analysis on Spheres and Balls
Feng Dai and Yuan Xu · 2013
Cited alongside, same era.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Cited alongside, same era.
On the spectral bias of neural networks
Nasim Rahaman, Aristide Baratin, Devansh Arpit, Felix Draxler, Min Lin, Fred Hamprecht, Yoshua Bengio, and Aaron Courville · 2019
Later among the works it cites.
Theory of the frequency principle for general deep neural networks
Tao Luo, Zheng Ma, Zhi-Qin John Xu, and Yaoyu Zhang · 2019
Later among the works it cites.
More data can hurt for linear regression: Sample-wise double descent
Preetum Nakkiran · 2019
Later among the works it cites.
A jamming transition from under-to over-parametrization affects generalization in deep learning
S Spigler, M Geiger, S d’Ascoli, L Sagun, G Biroli, and M Wyart · 2019
Later among the works it cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
C Zhang, S Bengio, M Hardt, B Recht, and O Vinyals · 2016
Cited alongside, same era.
Statistical mechanics of optimal convex inference in high dimensions
Madhu Advani and Surya Ganguli · 2016
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, and Nathan Srebro · 2017
Cited alongside, same era.
Gaussian process behaviour in wide deep neural networks
Alexander G de G Matthews, Jiri Hron, Mark Rowland, Richard E Turner, and Zoubin Ghahramani · 2018
Cited alongside, same era.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Cited alongside, same era.
Closest in time.
Kernel alignment risk estimator: Risk prediction from training data
Arthur Jacot, Berfin Simsek, Francesco Spadaro, Clement Hongler, and Franck Gabriel · 2020
Closest in time.
Just interpolate: Kernel “ridgeless” regression can generalize
Tengyuan Liang and Alexander Rakhlin · 2020
Closest in time.
Spectrum dependent learning curves in kernel regression and wide neural networks
Blake Bordelon, Abdulkadir Canatar, and Cengiz Pehlevan · 2020
Closest in time.
Implicit regularization of random feature models
Arthur Jacot, Berfin Simsek, Francesco Spadaro, Clement Hongler, and Franck Gabriel · 2020
Closest in time.
On the multiple descent of minimum-norm interpolants and restricted lower isometry of kernels
Tengyuan Liang, Alexander Rakhlin, and Xiyu Zhai · 2020
Closest in time.
A brief prehistory of double descent
Marco Loog, Tom Viering, Alexander Mey, Jesse H Krijthe, and David MJ Tax · 2020
Closest in time.
Double trouble in double descent : Bias and variance(s) in the lazy regime
Stéphane d’Ascoli, Maria Refinetti, Giulio Biroli, and Florent Krzakala · 2020
Closest in time.
Triple descent and the two kinds of overfitting: Where and why do they appear?
Stéphane d’Ascoli, Levent Sagun, and Giulio Biroli · 2020
Closest in time.
Neural tangents: Fast and easy infinite neural networks in python
Roman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee, Alexander A. Alemi, Jascha Sohl-Dickstein, and Samuel S. Schoenholz · 2020
Closest in time.
Tensor programs ii: Neural tangent kernel for any architecture
Greg Yang · 2020
Closest in time.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lénaïc Chizat and Francis Bach · 2020
Closest in time.
Asymptotic learning curves of kernel methods: empirical data versus teacher–student paradigm
Stefano Spigler, Mario Geiger, and Matthieu Wyart · 2020
Closest in time.
A random matrix analysis of random fourier features: beyond the gaussian kernel, a precise phase transition, and the corresponding double descent
Zhenyu Liao, Romain Couillet, and Michael W. Mahoney · 2020
Closest in time.
Towards understanding the spectral bias of deep learning
Yuan Cao, Zhiying Fang, Yue Wu, Ding-Xuan Zhou, and Quanquan Gu · 2020
Closest in time.
Generalization bounds for deep learning
Guillermo Valle-Pérez and Ard A. Louis · 2020
Closest in time.
Generalisation error in learning with random features and the hidden manifold model
Federica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mezard, and Lenka Zdeborova · 2020
Closest in time.
The neural tangent kernel in high dimensions: Triple descent and a multi-scale theory of generalization
Ben Adlam and Jeffrey Pennington · 2020
Closest in time.
Multiple descent: Design your own generalization curve
Lin Chen, Yifei Min, Mikhail Belkin, and Amin Karbasi · 2020
Closest in time.
Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the neural tangent kernel
Stanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani, Daniel M Roy, and Surya Ganguli · 2020
Closest in time.
{GS}hard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2021
Closest in time.
Optimal regularization can mitigate double descent
Preetum Nakkiran, Prayaag Venkat, Sham M. Kakade, and Tengyu Ma · 2021
Closest in time.
The recurrent neural tangent kernel
Sina Alemohammad, Zichao Wang, Randall Balestriero, and Richard Baraniuk · 2021
Closest in time.