Fetching the paper…
Reading the bibliography…
Modern machine learning often operates in the regime where the number of parameters is much higher than the number of data points, with zero training loss and yet good generalization, thereby contradicting the classical bias-variance trade-off.
Neural networks and the bias/variance dilemma
Stuart Geman, Elie Bienenstock, and René Doursat · 1992
Earlier work this paper cites.
Support-vector networks
Corinna Cortes and Vladimir Vapnik · 1995
Earlier work this paper cites.
The elements of statistical learning
Jerome Friedman, Trevor Hastie, and Robert Tibshirani · 2001
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
Optimal rates for the regularized least-squares algorithm
Andrea Caponnetto and Ernesto De Vito · 2007
Earlier work this paper cites.
Support vector machines
Ingo Steinwart and Andreas Christmann · 2008
Earlier work this paper cites.
Sharp analysis of low-rank kernel matrix approximations
Francis Bach · 2013
Earlier work this paper cites.
Steps toward deep kernel methods from infinite neural networks
Tamir Hazan and Tommi Jaakkola · 2015
Earlier work this paper cites.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Earlier work this paper cites.
On the equivalence between kernel quadrature rules and random feature expansions
Francis Bach · 2017
Earlier work this paper cites.
Random Fourier features for kernel ridge regression: Approximation bounds and statistical guarantees
Haim Avron, Michael Kapralov, Cameron Musco, Christopher Musco, Ameya Velingker, and Amir Zandieh · 2017
Earlier work this paper cites.
Generalization properties of learning with random features
Alessandro Rudi and Lorenzo Rosasco · 2017
Cited alongside, same era.
Sgd learns the conjugate kernel class of the network
Amit Daniely · 2017
Cited alongside, same era.
On some extensions of bernstein’s inequality for self-adjoint operators
Stanislav Minsker · 2017
Cited alongside, same era.
To understand deep learning we need to understand kernel learning
Mikhail Belkin, Siyuan Ma, and Soumik Mandal · 2018
Cited alongside, same era.
Just interpolate: Kernel" ridgeless" regression can generalize
Tengyuan Liang and Alexander Rakhlin · 2018
Cited alongside, same era.
Gaussian process behaviour in wide deep neural networks
AGG De Matthews, J Hron, M Rowland, RE Turner, and Z Ghahramani · 2018
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Later among the works it cites.
Two models of double descent for weak features
Mikhail Belkin, Daniel Hsu, and Ji Xu · 2019
Later among the works it cites.
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani · 2019
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and double descent curve
Song Mei and Andrea Montanari · 2019
Later among the works it cites.
Towards a unified analysis of random fourier features
Zhu Li, Jean-Francois Ton, Dino Oglic, and Dino Sejdinovic · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Bayesian deep convolutional networks with many channels are gaussian processes
Roman Novak, Lechao Xiao, Jaehoon Lee, Yasaman Bahri, Greg Yang, Jiri Hron, Daniel A Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Cited alongside, same era.
Deep convolutional networks as shallow gaussian processes
Adrià Garriga-Alonso, Carl Edward Rasmussen, and Laurence Aitchison · 2018
Cited alongside, same era.
Gaussian processes and kernel methods: A review on connections and equivalences
Motonobu Kanagawa, Philipp Hennig, Dino Sejdinovic, and Bharath K Sriperumbudur · 2018
Cited alongside, same era.
A mean field view of the landscape of two-layer neural networks
Song Mei, Andrea Montanari, and Phan-Minh Nguyen · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
High-dimensional probability: An introduction with applications in data science
Roman Vershynin · 2018
Cited alongside, same era.
Later among the works it cites.
Convergence types and rates in generic karhunen-loève expansions with applications to sample path properties
Ingo Steinwart · 2019
Later among the works it cites.
Gradient descent finds global minima of deep neural networks
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Later among the works it cites.
Zhenyu Liao, Romain Couillet, and Michael W Mahoney · 2020
Closest in time.
Benign overfitting in linear regression
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler · 2020
Closest in time.
The neural tangent kernel in high dimensions: Triple descent and a multi-scale theory of generalization
Ben Adlam and Jeffrey Pennington · 2020
Closest in time.
Taiji Suzuki · 2020
Closest in time.
Mean field analysis of neural networks: A central limit theorem
Justin Sirignano and Konstantinos Spiliopoulos · 2020
Closest in time.