Fetching the paper…
Reading the bibliography…
The remarkable practical success of deep learning has revealed some major surprises from a theoretical perspective.
On the multiple descent of minimum-norm interpolants and restricted lower isometry of kernels
Tengyuan Liang, Alexander Rakhlin, and Xiyu Zhai · 1908
Earlier work this paper cites.
The generalization error of random features regression: Precise asymptotics and double descent curve
Song Mei and Andrea Montanari · 1908
Earlier work this paper cites.
Asymptotics of wide networks from Feynman diagrams
Ethan Dyer and Guy Gur-Ari · 1909
Earlier work this paper cites.
Finite depth and width corrections to the neural tangent kernel
Boris Hanin and Mihai Nica · 1909
Earlier work this paper cites.
Beyond linearization: On quadratic and higher-order approximation of wide neural networks
Yu Bai and Jason D. Lee · 1910
Earlier work this paper cites.
Theoretical foundations of the potential function method in pattern recognition
MA Aizerman, E M Braverman, and LI Rozonoer · 1964
Earlier work this paper cites.
On estimating regression
Elizbar A Nadaraya · 1964
Earlier work this paper cites.
Smooth regression analysis
Geoffrey S Watson · 1964
Earlier work this paper cites.
Nearest neighbor pattern classification
Thomas Cover and Peter Hart · 1967
Earlier work this paper cites.
On the uniform convergence of relative frequencies of events to their probabilities
V. N. Vapnik and A. Ya. Chervonenkis · 1971
Earlier work this paper cites.
Theory of Pattern Recognition
V. N. Vapnik and A. Ya. Chervonenkis · 1974
Earlier work this paper cites.
The densest hemisphere problem
David S. Johnson and F. P. Preparata · 1978
Earlier work this paper cites.
Distribution-free inequalities for the deleted and holdout error estimates
Luc Devroye and Terry Wagner · 1979
Earlier work this paper cites.
Computers and Intractability: A Guide to the Theory of NP-Completeness
Michael R. Garey and David S. Johnson · 1979
Earlier work this paper cites.
Learnability and the Vapnik-Chervonenkis dimension
Anselm Blumer, A. Ehrenfeucht, David Haussler, and Manfred K. Warmuth · 1989
Earlier work this paper cites.
What size net gives valid generalization?
Eric B. Baum and David Haussler · 1989
Earlier work this paper cites.
A general lower bound on the number of examples needed for learning
A. Ehrenfeucht, David Haussler, Michael J. Kearns, and Leslie G. Valiant · 1989
Earlier work this paper cites.
Neural Network Design and the Complexity of Learning
J. S. Judd · 1990
Earlier work this paper cites.
Recognizing hand-printed letters and digits
Gale Martin and James Pittman · 1990
Earlier work this paper cites.
Empirical Processes: Theory and Applications
David Pollard · 1990
Earlier work this paper cites.
Estimating a regression function
Sara van de Geer · 1990
Earlier work this paper cites.
Probability in Banach Spaces: Isoperimetry and Processes
M. Ledoux and M. Talagrand · 1991
Earlier work this paper cites.
Training a 3-node neural network is NP-complete
Avrim Blum and Ronald L. Rivest · 1992
Earlier work this paper cites.
Decision theoretic generalizations of the PAC model for neural net and other learning applications
D. Haussler · 1992
Earlier work this paper cites.
Sharper bounds for Gaussian and empirical processes
M. Talagrand · 1994
Earlier work this paper cites.
Support-vector networks
Corinna Cortes and Vladimir Vapnik · 1995
Earlier work this paper cites.
Boosting decision trees
Harris Drucker and Corinna Cortes · 1995
Earlier work this paper cites.
On the complexity of training neural networks with continuous activation functions
Bhaskar DasGupta, Hava T. Siegelmann, and Eduardo D. Sontag · 1995
Earlier work this paper cites.
Uniform ratio limit theorems for empirical processes
David Pollard · 1995
Earlier work this paper cites.
Analysis of two simple heuristics on a random instance of k k -SAT
Alan Frieze and Stephen Suen · 1996
Earlier work this paper cites.
Efficient agnostic learning of neural networks with bounded fan-in
W. S. Lee, P. L. Bartlett, and R. C. Williamson · 1996
Earlier work this paper cites.
Bagging, boosting, and C4.5
J. R. Quinlan · 1996
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
R. Tibshirani · 1996
Earlier work this paper cites.
The analysis of a list-coloring algorithm on a random graph
Dimitris Achlioptas and Michael Molloy · 1997
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E. Schapire · 1997
Earlier work this paper cites.
Polynomial bounds for VC dimension of sigmoidal and general Pfaffian neural networks
Marek Karpinski and Angus J. Macintyre · 1997
Earlier work this paper cites.
Dimension-independent rates of approximation by neural networks
Věra Kurková · 1997
Earlier work this paper cites.
Lessons in neural network training: Overfitting may be harder than expected
Steve Lawrence, C. Lee Giles, and Ah Chung Tsoi · 1997
Earlier work this paper cites.
The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network
P. L. Bartlett · 1998
Earlier work this paper cites.
Almost linear VC dimension bounds for piecewise polynomial networks
P. L. Bartlett, V. Maiorov, and R. Meir · 1998
Earlier work this paper cites.
Arcing classifiers
Leo Breiman · 1998
Earlier work this paper cites.
The Hilbert kernel regression estimate
Luc Devroye, Laszlo Györfi, and Adam Krzyżak · 1998
Earlier work this paper cites.
Boosting the margin: A new explanation for the effectiveness of voting methods
Robert E Schapire, Yoav Freund, Peter Bartlett, and Wee Sun Lee · 1998
Earlier work this paper cites.
On the infeasibility of training neural networks with small squared errors
Van H. Vu · 1998
Earlier work this paper cites.
Neural Network Learning: Theoretical Foundations
Martin Anthony and Peter L. Bartlett · 1999
Earlier work this paper cites.
An inequality for uniform deviations of sample averages from their means
P. L. Bartlett and G. Lugosi · 1999
Earlier work this paper cites.
Rademacher processes and bounding the risk of function learning
V. I. Koltchinskii and D. Panchenko · 2000
Earlier work this paper cites.
Overfitting in neural nets: Backpropagation, conjugate gradient, and early stopping
Rich Caruana, Steve Lawrence, and C. Giles · 2001
Earlier work this paper cites.
The Elements of Statistical Learning
Jerome Friedman, Trevor Hastie, and Robert Tibshirani · 2001
Earlier work this paper cites.
Greedy function approximation: A gradient boosting machine
Jerome H. Friedman · 2001
Earlier work this paper cites.
Rademacher penalties and structural risk minimization
V. Koltchinskii · 2001
Earlier work this paper cites.
Bounds on rates of variable-basis and neural-network approximation
Vera Kurková and Marcello Sanguineti · 2001
Earlier work this paper cites.
Analytic study of double descent in binary classification: The impact of loss
Ganesh Ramachandra Kini and Christos Thrampoulidis · 2001
Earlier work this paper cites.
The concentration of measure phenomenon
Michel Ledoux · 2001
Earlier work this paper cites.
Using the Nyström method to speed up kernel machines
Christopher KI Williams and Matthias Seeger · 2001
Earlier work this paper cites.
Hardness results for neural network approximation problems
P. L. Bartlett and S. Ben-David · 2002
Earlier work this paper cites.
Model selection and error estimation
P. L. Bartlett, S. Boucheron, and G. Lugosi · 2002
Cited alongside, same era.
Stability and generalization
Olivier Bousquet and André Elisseeff · 2002
Cited alongside, same era.
Rademacher and Gaussian complexities: Risk bounds and structural results
P. L. Bartlett and S. Mendelson · 2002
Cited alongside, same era.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lénaïc Chizat and Francis Bach · 2002
Cited alongside, same era.
Comparison of worst case errors in linear and neural network approximation
Vera Kurková and Marcello Sanguineti · 2002
Cited alongside, same era.
Improving the sample complexity using global data
Shahar Mendelson · 2002
To understand deep learning we need to understand kernel learning
Mikhail Belkin, Siyuan Ma, and Soumik Mandal · 2018
Later among the works it cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Lénaïc Chizat and Francis Bach · 2018
Later among the works it cites.
Characterizing implicit bias in terms of optimization geometry
Suriya Gunasekar, Jason Lee, Daniel Soudry, and Nathan Srebro · 2018
Later among the works it cites.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason D Lee, Daniel Soudry, and Nati Srebro · 2018
Later among the works it cites.
Size-independent sample complexity of neural networks
Noah Golowich, Alexander Rakhlin, and Ohad Shamir · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A note on margin-based loss functions in classification
Y. Lin · 2004
Cited alongside, same era.
On the Bayes-risk consistency of regularized boosting methods
G. Lugosi and N. Vayatis · 2004
Cited alongside, same era.
Statistical behavior and consistency of classification methods based on convex risk minimization
Tong Zhang · 2004
Cited alongside, same era.
Local Rademacher complexities
Peter L. Bartlett, Olivier Bousquet, and Shahar Mendelson · 2005
Cited alongside, same era.
Boosting with early stopping: Convergence and consistency
Tong Zhang and Bin Yu · 2005
Cited alongside, same era.
Kernels as features: On kernels, margins, and low-dimensional mappings
Maria-Florina Balcan, Avrim Blum, and Santosh Vempala · 2006
Cited alongside, same era.
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Later among the works it cites.
Risk and parameter convergence of logistic regression
Ziwei Ji and Matus Telgarsky · 2018
Later among the works it cites.
A random matrix approach to neural networks
Cosme Louart, Zhenyu Liao, and Romain Couillet · 2018
Later among the works it cites.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2018
Later among the works it cites.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Later among the works it cites.
The spectrum of the Fisher information matrix of a single-hidden-layer neural network
Jeffrey Pennington and Pratik Worah · 2018
Later among the works it cites.
Grant Rotskoff and Eric Vanden-Eijnden · 2018
Later among the works it cites.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Later among the works it cites.
High-Dimensional Probability. An Introduction with Applications in Data Science
Roman Vershynin · 2018
Later among the works it cites.
A continuous-time view of early stopping for least squares regression
Alnur Ali, J. Zico Kolter, and Ryan J. Tibshirani · 2019
Later among the works it cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Later among the works it cites.
Nearly-tight VC-dimension and pseudodimension bounds for piecewise linear neural networks
Peter L. Bartlett, Nick Harvey, Christopher Liaw, and Abbas Mehrabian · 2019
Later among the works it cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Later among the works it cites.
Does data interpolation contradict statistical optimality?
Mikhail Belkin, Alexander Rakhlin, and Alexandre B Tsybakov · 2019
Later among the works it cites.
On lazy training in differentiable programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Later among the works it cites.
Gradient descent finds global minima of deep neural networks
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Later among the works it cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S. Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Later among the works it cites.
The spectral norm of random inner-product kernel matrices
Zhou Fan and Andrea Montanari · 2019
Later among the works it cites.
Modelling the influence of data structure on learning in neural networks
Sebastian Goldt, Marc Mézard, Florent Krzakala, and Lenka Zdeborová · 2019
Later among the works it cites.
Jamming transition as a paradigm to understand the loss landscape of deep neural networks
Mario Geiger, Stefano Spigler, Stéphane d’Ascoli, Levent Sagun, Marco Baity-Jesi, Giulio Biroli, and Matthieu Wyart · 2019
Later among the works it cites.
Gaussian Estimation: Sequence and Wavelet Models
Iain M. Johnstone · 2019
Later among the works it cites.
A refined primal-dual analysis of the implicit bias
Ziwei Ji and Matus Telgarsky · 2019
Later among the works it cites.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Later among the works it cites.
Andrea Montanari, Feng Ruan, Youngtak Sohn, and Jun Yan · 2019
Later among the works it cites.
Uniform convergence may be unable to explain generalization in deep learning
Vaishnavh Nagarajan and J. Zico Kolter · 2019
Later among the works it cites.
Convergence of gradient descent on separable data
Mor Shpigel Nacson, Jason Lee, Suriya Gunasekar, Pedro Henrique Pamplona Savarese, Nathan Srebro, and Daniel Soudry · 2019
Later among the works it cites.
Overparameterized nonlinear learning: Gradient descent takes the shortest path?
Samet Oymak and Mahdi Soltanolkotabi · 2019
Later among the works it cites.
Consistency of interpolation with Laplace kernels is a high-dimensional phenomenon
Alexander Rakhlin and Xiyu Zhai · 2019
Later among the works it cites.
A jamming transition from under-to over-parametrization affects generalization in deep learning
Stefano Spigler, Mario Geiger, Stéphane d’Ascoli, Levent Sagun, Giulio Biroli, and Matthieu Wyart · 2019
Later among the works it cites.
Failures of model-dependent generalization bounds for least-norm interpolation
Peter L. Bartlett and Philip M. Long · 2020
Later among the works it cites.
Benign overfitting in linear regression
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler · 2020
Later among the works it cites.
Does learning require memorization? A short tale about a long tail
Vitaly Feldman · 2020
Later among the works it cites.
Linearized two-layers neural networks in high dimension
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2020
Later among the works it cites.
When do neural networks outperform kernel methods?
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2020
Later among the works it cites.
The gaussian equivalence of generative models for learning with two-layer neural networks
Sebastian Goldt, Galen Reeves, Marc Mézard, Florent Krzakala, and Lenka Zdeborová · 2020
Later among the works it cites.
Universality laws for high-dimensional learning with random features
Hong Hu and Yue M Lu · 2020
Later among the works it cites.
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani · 2020
Later among the works it cites.
Dynamics of deep neural networks and neural tangent hierarchy
Jiaoyang Huang and Horng-Tzer Yau · 2020
Later among the works it cites.
Personal communication
Tengyuan Liang, 2020 · 2020
Later among the works it cites.
Just interpolate: Kernel “ridgeless” regression can generalize
Tengyuan Liang and Alexander Rakhlin · 2020
Later among the works it cites.
On the linearity of large non-linear models: when and why the tangent kernel is constant
Chaoyue Liu, Libin Zhu, and Mikhail Belkin · 2020
Later among the works it cites.
Extending the scope of the small-ball method
Shahar Mendelson · 2020
Later among the works it cites.
Andrea Montanari and Yiqiao Zhong · 2020
Later among the works it cites.
A rigorous framework for the mean field limit of multilayer neural networks
Phan-Minh Nguyen and Huy Tuan Pham · 2020
Later among the works it cites.
Towards moderate overparameterization: global convergence guarantees for training shallow neural networks
Samet Oymak and Mahdi Soltanolkotabi · 2020
Later among the works it cites.
Mean field analysis of neural networks: A law of large numbers
Justin Sirignano and Konstantinos Spiliopoulos · 2020
Later among the works it cites.
Benign overfitting in ridge regression
Alexander Tsigler and Peter L Bartlett · 2020
Later among the works it cites.
Fundamental limits of ridge-regularized empirical risk minimization in high dimensions
Hossein Taheri, Ramtin Pedarsani, and Christos Thrampoulidis · 2020
Later among the works it cites.
Gradient descent optimizes over-parameterized deep ReLU networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2020
Later among the works it cites.
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2021
Closest in time.