Fetching the paper…
Reading the bibliography…
In the context of neural network models, overparametrization refers to the phenomena whereby these models appear to generalize well on the unseen data, even though the number of parameters significantly exceeds the sample sizes, and the model perfectly fits the in-training data.
Issai Schur, Bemerkungen zur theorie der beschränkten bilinearformen mit unendlich vielen veränderlichen. , Journal für die reine und Angewandte Mathematik 140
1911
Earlier work this paper cites.
Donald J Newman et al., Rational approximation to | x | |x| . , The Michigan Mathematical Journal 11
1964
Earlier work this paper cites.
Walter Rudin et al., Principles of mathematical analysis , vol. 3, McGraw-hill New York, 1964
1964
Earlier work this paper cites.
Gabor Szegö, Orthogonal polynomials, vol. 23 , American Mathematical Society Colloquium Publications, 1975
1975
Earlier work this paper cites.
Wassily Hoeffding, Probability inequalities for sums of bounded random variables , The Collected Works of Wassily Hoeffding, Springer, 1994, pp. 409–426
1994
Earlier work this paper cites.
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner, Gradient-based learning applied to document recognition , Proceedings of the IEEE, 1998, pp. 2278–2324
1998
Earlier work this paper cites.
Yann LeCun, Corinna Cortes, and Christopher JC Burges, The mnist database of handwritten digits, 1998 , URL http://yann. lecun. com/exdb/mnist 10
1998
Earlier work this paper cites.
Richard Caron and Tim Traynor, The zero set of a polynomial , WSMR Report (2005), 05–02
2005
Earlier work this paper cites.
Vladimir Nikiforov, Revisiting schur’s bound on the largest singular value , arXiv preprint math/0702722 (2007)
2007
Earlier work this paper cites.
Katerina Hlavácková-Schindler, A new lower bound for the minimal singular value for real non-singular matrices by a matrix norm and determinant , Applied Mathematical Sciences 4
2010
Earlier work this paper cites.
2010
Earlier work this paper cites.
Roger A Horn and Charles R Johnson, Matrix analysis , Cambridge University Press, 2012
2012
Cited alongside, same era.
Li Wan, Matthew Zeiler, Sixin Zhang, Yann Le Cun, and Rob Fergus, Regularization of neural networks using dropconnect , Proceedings of the 30th International Conference on Machine Learning, Proceedings of Machine Learning Research, vol. 28:3, Jun 2013, pp. 1058–1066
2013
Cited alongside, same era.
Shai Shalev-Shwartz and Shai Ben-David, Understanding machine learning: From theory to algorithms , Cambridge university press, 2014
2014
Cited alongside, same era.
2015
Cited alongside, same era.
Suriya Gunasekar, Jason D Lee, Daniel Soudry, and Nati Srebro, Implicit bias of gradient descent on linear convolutional networks , Advances in Neural Information Processing Systems, 2018, pp. 9461–9471
2018
Later among the works it cites.
Arthur Jacot, Franck Gabriel, and Clément Hongler, Neural tangent kernel: Convergence and generalization in neural networks , Advances in neural information processing systems, 2018, pp. 8571–8580
2018
Later among the works it cites.
Yuanzhi Li and Yingyu Liang, Learning overparameterized neural networks via stochastic gradient descent on structured data , Advances in Neural Information Processing Systems, 2018, pp. 8157–8166
2018
Later among the works it cites.
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Russ R Salakhutdinov, and Ruosong Wang, On exact computation with an infinitely wide neural net , Advances in Neural Information Processing Systems, 2019, pp. 8139–8148
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
2017
Cited alongside, same era.
Matus Telgarsky, Neural networks and rational functions , Proceedings of the 34th International Conference on Machine Learning-Volume 70, JMLR. org, 2017, pp. 3387–3393
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Later among the works it cites.
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal, Reconciling modern machine-learning practice and the classical bias–variance trade-off , Proceedings of the National Academy of Sciences 116
2019
Later among the works it cites.
Sebastian Goldt, Madhu Advani, Andrew M Saxe, Florent Krzakala, and Lenka Zdeborová, Dynamics of stochastic gradient descent for two-layer neural networks in the teacher-student setup , Advances in Neural Information Processing Systems, 2019, pp. 6979–6989
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
Roman Vershynin, Concentration inequalities for random tensors , 2019
2019
Later among the works it cites.