Fetching the paper…
Reading the bibliography…
Modern neural networks are often operated in a strongly overparametrized regime: they comprise so many parameters that they can interpolate the training set, even if actual labels are replaced by purely random ones.
1908
Earlier work this paper cites.
Isaac J. Schoenberg, Positive definite functions on spheres , Duke Math. J. 9
1942
Earlier work this paper cites.
Thomas M Cover, Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition , IEEE transactions on electronic computers (1965), no. 3, 326–334
1965
Earlier work this paper cites.
Chandler Davis and W. M. Kahan, The rotation of eigenvectors by a perturbation. iii , SIAM Journal on Numerical Analysis 7
1970
Earlier work this paper cites.
Eric B Baum, On the capabilities of multilayer perceptrons , Journal of Complexity 4
1988
Earlier work this paper cites.
Akito Sakurai, n-h-1 networks store no less n*h+1 examples, but sometimes no more , [Proceedings 1992] IJCNN International Joint Conference on Neural Networks, vol. 3, 1992, pp. 936–941
1992
Earlier work this paper cites.
Adam Kowalczyk, Counting function theorem for multi-layer networks , Advances in neural information processing systems, 1994, pp. 375–382
1994
Earlier work this paper cites.
Ali Rahimi and Benjamin Recht, Random features for large-scale kernel machines , Advances in neural information processing systems, 2008, pp. 1177–1184
2008
Earlier work this paper cites.
2009
Earlier work this paper cites.
2010
Earlier work this paper cites.
Joel Tropp et al., Freedman’s inequality for matrix martingales , Electronic Communications in Probability 16
2011
Earlier work this paper cites.
Kendall Atkinson and Weimin Han, Spherical harmonics and approximations on the unit sphere: an introduction , vol. 2044, Springer Science, 2012
2012
Earlier work this paper cites.
Efthimiou Costas and Frye Christopher, Spherical harmonics in p p dimensions , World Scientific, 2014
2014
Earlier work this paper cites.
Ramon van Handel, Probability in high dimension , Tech. report, PRINCETON UNIV NJ, 2014
2014
Earlier work this paper cites.
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro, In search of the real inductive bias: On the role of implicit regularization in deep learning. , ICLR (Workshop), 2015
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
Francis Bach, Breaking the curse of dimensionality with convex neural networks , The Journal of Machine Learning Research 18
2017
Earlier work this paper cites.
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh, Gradient descent provably optimizes over-parameterized neural networks , International Conference on Learning Representations, 2018
2018
Cited alongside, same era.
Arthur Jacot, Franck Gabriel, and Clément Hongler, Neural tangent kernel: Convergence and generalization in neural networks , Advances in neural information processing systems, 2018, pp. 8571–8580
2018
Cited alongside, same era.
Mahdi Soltanolkotabi, Adel Javanmard, and Jason D Lee, Theoretical insights into the optimization landscape of over-parameterized shallow neural networks , IEEE Transactions on Information Theory 65
2018
Cited alongside, same era.
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, and Ruosong Wang, Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks , International Conference on Machine Learning, PMLR, 2019, pp. 322–332
2019
Cited alongside, same era.
2019
Later among the works it cites.
2019
Later among the works it cites.
Chulhee Yun, Suvrit Sra, and Ali Jadbabaie, Small relu networks are powerful memorizers: a tight analysis of memorization capacity , Advances in Neural Information Processing Systems, 2019, pp. 15558–15569
2019
Later among the works it cites.
2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang, Learning and generalization in overparameterized neural networks, going beyond two layers , Advances in Neural Information Processing Systems 32
2019
Cited alongside, same era.
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song, A convergence theory for deep learning via over-parameterization , International Conference on Machine Learning, 2019, pp. 242–252
2019
Cited alongside, same era.
Peter L Bartlett, Nick Harvey, Christopher Liaw, and Abbas Mehrabian, Nearly-tight vc-dimension and pseudodimension bounds for piecewise linear neural networks. , J. Mach. Learn. Res. 20
2019
Cited alongside, same era.
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal, Reconciling modern machine-learning practice and the classical bias–variance trade-off , Proceedings of the National Academy of Sciences 116
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Yuan Cao and Quanquan Gu, Generalization bounds of stochastic gradient descent for wide and deep neural networks , Advances in Neural Information Processing Systems 32
2019
Cited alongside, same era.
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler, Benign overfitting in linear regression , Proceedings of the National Academy of Sciences (2020)
2020
Closest in time.
2020
Closest in time.
Tengyuan Liang and Alexander Rakhlin, Just interpolate: Kernel “ridgeless” regression can generalize , Annals of Statistics 48
2020
Closest in time.
Chaoyue Liu, Libin Zhu, and Mikhail Belkin, On the linearity of large non-linear models: when and why the tangent kernel is constant , Advances in Neural Information Processing Systems (H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, eds.), vol. 33, Curran Associates, Inc., 2020, pp. 15954–15964
2020
Closest in time.
Samet Oymak and Mahdi Soltanolkotabi, Towards moderate overparameterization: global convergence guarantees for training shallow neural networks , IEEE Journal on Selected Areas in Information Theory (2020)
2020
Closest in time.
E Weinan, Ma Chao, and Wu Lei, A comparative analysis of optimization and generalization properties oftwo-layer neural network and random feature models under gradient descent dynamics , Science China Mathematics 63
2020
Closest in time.
2020
Closest in time.
Peter L. Bartlett, Andrea Montanari, and Alexander Rakhlin, Deep learning: a statistical viewpoint , Acta Numerica 30
2021
Closest in time.
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari, Linearized two-layers neural networks in high dimension , The Annals of Statistics 49
2021
Closest in time.
Song Mei and Andrea Montanari, The generalization error of random features regression: Precise asymptotics and double descent curve , Communications in Pure and Applied Mathematics (2021)
2021
Closest in time.
2021
Closest in time.
Sejun Park, Jaeho Lee, Chulhee Yun, and Jinwoo Shin, Provable memorization via deep neural networks using sub-linear parameters , Conference on Learning Theory, PMLR, 2021, pp. 3627–3661
2021
Closest in time.