Fetching the paper…
Reading the bibliography…
Why do deep neural networks (DNNs) benefit from very high dimensional parameter spaces? Their huge parameter complexities vs stunning performance in practice is all the more intriguing and not explainable using the standard theory of model selection for regular models.
Geometry of lightlike hypersurfaces of a statistical manifold, 2019
Oguzhan Bahadir and Mukut Mani Tripathi · 1901
Earlier work this paper cites.
Negative eigenvalues of the Hessian in deep neural networks
Guillaume Alain, Nicolas Le Roux, and Pierre-Antoine Manzagol · 1902
Earlier work this paper cites.
On the geometry of lightlike submanifolds of indefinite statistical manifolds, 2019
Varun Jain, Amrinder Pal Singh, and Rakesh Kumar · 1903
Earlier work this paper cites.
Spaces of statistical parameters
Harold Hotelling · 1930
Earlier work this paper cites.
Sur la notion de la moyenne
Andreĭ Nikolaevich Kolmogorov · 1930
Earlier work this paper cites.
Über eine Klasse der Mittelwerte
Mitio Nagumo · 1930
Earlier work this paper cites.
Information and the accuracy attainable in the estimation of statistical parameters
Calyampudi Radhakrishna Rao · 1945
Earlier work this paper cites.
An information measure for classification
Christopher Stewart Wallace and D. M. Boulton · 1968
Earlier work this paper cites.
A new look at the statistical model identification
Hirotugu Akaike · 1974
Earlier work this paper cites.
Modeling by shortest data description
Jorma Rissanen · 1978
Earlier work this paper cites.
Estimating the dimension of a model
Gideon Schwarz · 1978
Earlier work this paper cites.
Statistical manifolds
Stefan L Lauritzen · 1987
Earlier work this paper cites.
Universal sequential coding of single messages
Y. M. Shtarkov · 1987
Earlier work this paper cites.
Schaum’s outline of theory and problems of tensor calculus
David C Kay · 1988
Earlier work this paper cites.
Minimum complexity density estimation
A.R. Barron and T.M. Cover · 1991
Earlier work this paper cites.
Bayesian methods for adaptive models
David J.C. MacKay · 1992
Earlier work this paper cites.
Information and the accuracy attainable in the estimation of statistical parameters
Calyampudi Radhakrishna Rao · 1992
Earlier work this paper cites.
Network information criterion-determining the number of hidden units for an artificial neural network model
Noboru Murata, Shuji Yoshizawa, and Shun-ichi Amari · 1994
Earlier work this paper cites.
Affine differential geometry: geometry of affine immersions
Katsumi Nomizu, Nomizu Katsumi, and Takeshi Sasaki · 1994
Earlier work this paper cites.
Lightlike Submanifolds of Semi-Riemannian Manifolds and Applications
Krishan Duggal and Aurel Bejancu · 1996
Earlier work this paper cites.
Singular Semi-Riemannian Geometry
D.N. Kupeli · 1996
Earlier work this paper cites.
Fisher information and stochastic complexity
Jorma Rissanen · 1996
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
The minimum description length principle in coding and modeling
A. Barron, J. Rissanen, and Bin Yu · 1998
Earlier work this paper cites.
Counting probability distributions: Differential geometry and model selection
In Jae Myung, Vijay Balasubramanian, and Mark A. Pitt · 2000
Earlier work this paper cites.
Strong optimality of the normalized ml models as universal codes and information in data
J. Rissanen · 2001
Earlier work this paper cites.
PAC-MDL bounds
A. Blum and J. Langford · 2003
Earlier work this paper cites.
Singularities in mixture models and upper bounds of stochastic complexity
Keisuke Yamazaki and Sumio Watanabe · 2003
Earlier work this paper cites.
MDL, Bayesian inference and the geometry of the space of probability distributions
Vijay Balasubramanian · 2005
Cited alongside, same era.
α \alpha -parallel prior and its properties
Junnichi Takeuchi and S-I Amari · 2005
Cited alongside, same era.
Information-theoretic upper and lower bounds for statistical estimation
Tong Zhang · 2006
Cited alongside, same era.
The rank of a random matrix
Xinlong Feng and Zhinan Zhang · 2007
Cited alongside, same era.
The Minimum Description Length Principle
Peter D. Grünwald · 2007
Cited alongside, same era.
Dynamics of learning near singularities in layered networks
Haikun Wei, Jun Zhang, Florent Cousseau, Tomoko Ozeki, and Shun-ichi Amari · 2008
Cited alongside, same era.
Dynamics of learning in MLP: Natural gradient and singularity revisited
Shun-ichi Amari, Tomoko Ozeki, Ryo Karakida, Yuki Yoshida, and Masato Okada · 2018
Later among the works it cites.
The description length of deep learning models
Léonard Blier and Yann Ollivier · 2018
Later among the works it cites.
Model compression and acceleration for deep neural networks: The principles, progress, and challenges
Yu Cheng, Duo Wang, Pan Zhou, and Tao Zhang · 2018
Later among the works it cites.
Measuring the intrinsic dimension of objective landscapes
Chunyuan Li, Heerad Farkhoor, Rosanne Liu, and Jason Yosinski · 2018
Later among the works it cites.
Skip connections eliminate singularities
A Emin Orhan and Xaq Pitkow · 2018
Later among the works it cites.
The emergence of spectral universality in deep networks
Jeffrey Pennington, Samuel Schoenholz, and Surya Ganguli · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Algebraic Geometry and Statistical Learning Theory
Sumio Watanabe · 2009
Cited alongside, same era.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Cited alongside, same era.
Deep sparse rectifier neural networks
Xavier Glorot, Antoine Bordes, and Yoshua Bengio · 2011
Cited alongside, same era.
Large deviation theory for non-regular location shift family
Masahito Hayashi · 2011
Cited alongside, same era.
A note on insufficiency and the preservation of Fisher information
David Pollard · 2013
Cited alongside, same era.
Geometric modeling in probability and statistics
Ovidiu Calin and Constantin Udrişte · 2014
Cited alongside, same era.
Later among the works it cites.
The spectrum of the Fisher information matrix of a single-hidden-layer neural network
Jeffrey Pennington and Pratik Worah · 2018
Later among the works it cites.
Empirical analysis of the Hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V. Ugur Guney, Yann Dauphin, and Leon Bottou · 2018
Later among the works it cites.
Weight agnostic neural networks
Adam Gaier and David Ha · 2019
Closest in time.
Minimum description length revisited
Peter Grünwald and Teemu Roos · 2019
Closest in time.
A tight excess risk bound via a unified PAC-Bayesian–Rademacher–Shtarkov–MDL complexity
Peter D. Grünwald and Nishant A. Mehta · 2019
Closest in time.
Universal statistics of Fisher information in deep neural networks: Mean field approach
Ryo Karakida, Shotaro Akaho, and Shun-ichi Amari · 2019
Closest in time.
Limitations of the empirical Fisher approximation for natural gradient descent
Frederik Kunstner, Philipp Hennig, and Lukas Balles · 2019
Closest in time.
Fisher-Rao metric, geometry, and complexity of neural networks
Tengyuan Liang, Tomaso Poggio, Alexander Rakhlin, and James Stokes · 2019
Closest in time.
Deep learning generalizes because the parameter-function map is biased towards simple functions
Guillermo Valle-Pérez, Chico Q. Camargo, and Ard A. Louis · 2019
Closest in time.
Statistical mechanical analysis of learning dynamics of two-layer perceptron with multiple output units
Yuki Yoshida, Ryo Karakida, Masato Okada, and Shun-ichi Amari · 2019
Closest in time.
Deep learning architectures
Ovidiu Calin · 2020
Closest in time.
Weyl prior and Bayesian statistics
Ruichao Jiang, Javad Tavakoli, and Yiqiang Zhao · 2020
Closest in time.
New insights and perspectives on the natural gradient method
James Martens · 2020
Closest in time.
Traces of class/cross-class structure pervade deep learning spectra
Vardan Papyan · 2020
Closest in time.
Towards modeling and resolving singular parameter spaces using stratifolds
Pascal Mattia Esser and Frank Nielsen · 2021
Closest in time.
The spectrum of Fisher information of deep networks achieving dynamical isometry
Tomohiro Hayase and Ryo Karakida · 2021
Closest in time.
Pathological Spectra of the Fisher Information Metric and Its Variants in Deep Neural Networks
Ryo Karakida, Shotaro Akaho, and Shun-ichi Amari · 2021
Closest in time.
A unified formulation of k k -Means, fuzzy c c -Means and Gaussian mixture model by the Kolmogorov–Nagumo average
Osamu Komori and Shinto Eguchi · 2021
Closest in time.
The dually flat structure for singular models
Naomichi Nakajima and Toru Ohmoto · 2021
Closest in time.
On the variance of the Fisher information for deep learning
Alexander Soen and Ke Sun · 2021
Closest in time.
Simplifying momentum-based positive-definite submanifold optimization with applications to deep learning
Wu Lin, Valentin Duruisseaux, Melvin Leok, Frank Nielsen, Mohammad Emtiyaz Khan, and Mark Schmidt · 2023
Closest in time.