Fetching the paper…
Reading the bibliography…
We study in this paper lower bounds for the generalization error of models derived from multi-layer neural networks, in the regime where the size of the layers is commensurate with the number of samples in the training data.
A contribution to the theory of statistical estimation
Harald Cramér · 1946
Earlier work this paper cites.
Minimum variance and the estimation of several parameters
C Radhakrishna Rao · 1947
Earlier work this paper cites.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Distribution of eigenvalues for some sets of random matrices
Vladimir A Marčenko and Leonid Andreevich Pastur · 1967
Earlier work this paper cites.
A modified Cramér-Rao bound and its applications
R Miller and Chow Chang · 1978
Earlier work this paper cites.
Some classes of global cramér-rao bounds
Ben-Zion Bobrovsky, E Mayer-Wolf, and M Zakai · 1987
Earlier work this paper cites.
Statistical mechanics of learning from examples
Hyunjune Sebastian Seung, Haim Sompolinsky, and Naftali Tishby · 1992
Earlier work this paper cites.
Applications of the van Trees inequality: a Bayesian Cramér-Rao bound
Richard D Gill, Boris Y Levit, et al · 1995
Earlier work this paper cites.
Strong convergence of the empirical distribution of eigenvalues of large dimensional random matrices
Jack W Silverstein · 1995
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
Posterior cramér-rao bounds for discrete-time nonlinear filtering
Petr Tichavsky, Carlos H Muravchik, and Arye Nehorai · 1998
Earlier work this paper cites.
Pac-bayesian model averaging
David A McAllester · 1999
Earlier work this paper cites.
Statistical mechanics of learning
Andreas Engel and Christian Van den Broeck · 2001
Earlier work this paper cites.
Detection, estimation, and modulation theory, part I: detection, estimation, and linear modulation theory
Harry L Van Trees · 2004
Earlier work this paper cites.
Theory of point estimation
Erich L Lehmann and George Casella · 2006
Earlier work this paper cites.
Rethinking biased estimation: Improving maximum likelihood and the Cramér-Rao bound
Yonina C Eldar · 2008
Earlier work this paper cites.
A lower bound on the bayesian mse based on the optimal bias function
Zvika Ben-Haim and Yonina C Eldar · 2009
Earlier work this paper cites.
The spectrum of kernel random matrices
Noureddine El Karoui et al · 2010
Cited alongside, same era.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Geoffrey Hinton, Li Deng, Dong Yu, George E Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N Sainath, et al · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Cited alongside, same era.
Revisiting the van trees inequality in the spirit of hajek and le cam
Elisabeth Gassiat, David Pollard, and Gilles Stoltz · 2013
Cited alongside, same era.
Deep learning in neural networks: An overview
Jürgen Schmidhuber · 2015
Cited alongside, same era.
On the uniform convergence of relative frequencies of events to their probabilities
A random matrix approach to neural networks
Cosme Louart, Zhenyu Liao, Romain Couillet, et al · 2018
Later among the works it cites.
The spectrum of the fisher information matrix of a single-hidden-layer neural network
Jeffrey Pennington and Pratik Worah · 2018
Later among the works it cites.
Nearly-tight VC-dimension and pseudodimension bounds for piecewise linear neural networks
Peter L Bartlett, Nick Harvey, Christopher Liaw, and Abbas Mehrabian · 2019
Later among the works it cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Later among the works it cites.
Two models of double descent for weak features
Mikhail Belkin, Daniel Hsu, and Ji Xu · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vladimir N Vapnik and A Ya Chervonenkis · 2015
Cited alongside, same era.
Ridge regression and asymptotic minimax estimation over spheres of growing dimension
Lee H Dicker et al · 2016
Cited alongside, same era.
Deep learning
Ian Goodfellow, Yoshua Bengio, Aaron Courville, and Yoshua Bengio · 2016
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Cited alongside, same era.
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Cited alongside, same era.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nati Srebro · 2017
Cited alongside, same era.
A pac-bayesian approach to spectrally-normalized margin bounds for neural networks
Behnam Neyshabur, Srinadh Bhojanapalli, and Nathan Srebro · 2017
Cited alongside, same era.
Lucas Benigni and Sandrine Péché · 2019
Later among the works it cites.
Linearized two-layers neural networks in high dimension
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Later among the works it cites.
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani · 2019
Later among the works it cites.
Fantastic generalization measures and where to find them
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 2019
Later among the works it cites.
Fisher-rao metric, geometry, and complexity of neural networks
Tengyuan Liang, Tomaso Poggio, Alexander Rakhlin, and James Stokes · 2019
Later among the works it cites.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and double descent curve
Song Mei and Andrea Montanari · 2019
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2019
Later among the works it cites.
Benign overfitting in linear regression
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler · 2020
Later among the works it cites.
Spectra of the conjugate kernel and neural tangent kernel for linear-width neural networks
Zhou Fan and Zhichao Wang · 2020
Later among the works it cites.
Harmless interpolation of noisy data in regression
Vidya Muthukumar, Kailas Vodrahalli, Vignesh Subramanian, and Anant Sahai · 2020
Later among the works it cites.