Fetching the paper…
Reading the bibliography…
In singular models, the optimal set of parameters forms an analytic set with singularities and classical statistical inference cannot be applied to such models.
Information matrices and generalization
Valentin Thomas, Fabian Pedregosa, Bart van Merriënboer, Pierre-Antoine Mangazol, Yoshua Bengio, and Nicolas Le Roux · 1906
Earlier work this paper cites.
Keeping the neural networks simple by minimizing the description length of the weights
Geoffrey E Hinton and Drew Van Camp · 1993
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Statistical inference, Occam’s razor and statistical mechanics on the space of probability distributions
Vijay Balasubramanian · 1997
Earlier work this paper cites.
Learning and inference in hierarchical models with singularities
Shun-ichi Amari, Tomoko Ozeki, and Hyeyoung Park · 2003
Earlier work this paper cites.
Introduction to statistical learning theory
Olivier Bousquet, Stéphane Boucheron, and Gábor Lugosi · 2003
Earlier work this paper cites.
Almost All Learning Machines are Singular
Sumio Watanabe · 2007
Earlier work this paper cites.
Variational Bayes Solution of Linear Neural Networks and Its Generalization Performance
Shinichi Nakajima and Sumio Watanabe · 2007
Earlier work this paper cites.
Algebraic Geometry and Statistical Learning Theory
Sumio Watanabe · 2009
Earlier work this paper cites.
A Widely Applicable Bayesian Information Criterion
Sumio Watanabe · 2013
Earlier work this paper cites.
The No-U-Turn sampler: adaptively setting path lengths in hamiltonian monte carlo
Matthew D Hoffman and Andrew Gelman · 2014
Earlier work this paper cites.
Norm-based capacity control in neural networks
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Earlier work this paper cites.
Gaussian error linear units (GELUs)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Singularity of the Hessian in deep learning
Levent Sagun, Léon Bottou, and Yann LeCun · 2016
Earlier work this paper cites.
An introduction to differentiable manifolds and Riemannian geometry
William M Boothby · 2016
Earlier work this paper cites.
A Bayesian perspective on generalization and stochastic gradient descent
Samuel L Smith and Quoc V Le · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Cited alongside, same era.
SGD learns the conjugate kernel class of the network
Amit Daniely · 2017
Cited alongside, same era.
Entropy-SGD: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2017
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Later among the works it cites.
Mathematical Theory of Bayesian Statistics
Sumio Watanabe · 2018
Later among the works it cites.
Theory of deep learning III: explaining the non-overfitting puzzle
Tomaso A. Poggio, Kenji Kawaguchi, Qianli Liao, Brando Miranda, Lorenzo Rosasco, Xavier Boix, Jack Hidary, and Hrushikesh Mhaskar · 2018
Later among the works it cites.
Stochastic natural gradient descent draws posterior samples in function space
Samuel L Smith, Daniel Duckworth, Semon Rezchikov, Quoc V Le, and Jascha Sohl-Dickstein · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stanislaw Jastrzebski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Cited alongside, same era.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nati Srebro · 2017
Cited alongside, same era.
Swish: a self-gated activation function
Prajit Ramachandran, Barret Zoph, and Quoc V Le · 2017
Cited alongside, same era.
Fractional Langevin Monte Carlo: exploring Levy driven stochastic differential equations for Markov Chain Monte Carlo
Umut ŞimŠekli · 2017
Cited alongside, same era.
Stochastic gradient descent as approximate Bayesian inference
Stephan Mandt, Matthew D Hoffman, and David M Blei · 2017
Cited alongside, same era.
Deep learning scaling is predictable, empirically
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory F. Diamos, Heewoo Jun, Hassan Kianinejad, Md. Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou · 2017
Cited alongside, same era.
Energy-entropy competition and the effectiveness of stochastic gradient descent in machine learning
Yao Zhang, Andrew M. Saxe, Madhu S. Advani, and Alpha A. Lee · 2018
Cited alongside, same era.
The spectrum of the Fisher information matrix of a single-hidden-layer neural network
Jeffrey Pennington and Pratik Worah · 2018
Later among the works it cites.
Towards understanding the role of over-parametrization in generalization of neural networks
Behnam Neyshabur and Zhiyuan Li · 2019
Later among the works it cites.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Later among the works it cites.
On the power and limitations of random features for understanding neural networks
Gilad Yehudai and Ohad Shamir · 2019
Later among the works it cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Yuan Cao and Quanquan Gu · 2019
Later among the works it cites.
Rethinking parameter counting in deep models: Effective dimensionality revisited
Wesley J Maddox, Gregory Benton, and Andrew Gordon Wilson · 2020
Closest in time.
Being Bayesian, even just a bit, fixes overconfidence in ReLU networks
Agustinus Kristiadi, Matthias Hein, and Philipp Hennig · 2020
Closest in time.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Closest in time.
Functional vs. parametric equivalence of ReLU networks
Mary Phuong and Christoph H. Lampert · 2020
Closest in time.