Fetching the paper…
Reading the bibliography…
The Local Learning Coefficient (LLC) is introduced as a novel complexity measure for deep neural networks (DNNs).
Estimating real log canonical thresholds
Imai, T. (2019a) · 1906
Earlier work this paper cites.
On the overestimation of widely applicable Bayesian information criterion
Imai, T. (2019b) · 1908
Earlier work this paper cites.
Resolution of Singularities of an Algebraic Variety Over a Field of Characteristic Zero: I
Hironaka, H. (1964) · 1964
Earlier work this paper cites.
On the uniform convergence of relative frequencies of events to their probabilities
Vapnik, V. N. and Chervonenkis, A. Y. (1971) · 1971
Earlier work this paper cites.
Uniqueness of the weights for minimal feedforward nets with a given input-output map
Sussmann, H. J. (1992) · 1992
Earlier work this paper cites.
Reconstructing a neural net from its output
Fefferman, C. (1994) · 1994
Earlier work this paper cites.
Functionally equivalent feedforward neural networks
Kůrková, V. and Kainen, P. C. (1994) · 1994
Earlier work this paper cites.
A regularity condition of the information matrix of a multilayer perceptron network
Fukumizu, K. (1996) · 1996
Earlier work this paper cites.
Statistical inference, Occam’s razor, and statistical mechanics on the space of probability distributions
Balasubramanian, V. (1997) · 1997
Earlier work this paper cites.
Flat Minima
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Optimal scaling of discrete approximations to Langevin diffusions
Roberts, G. O. and Rosenthal, J. S. (1998) · 1998
Earlier work this paper cites.
Adaptive proposal distribution for random walk Metropolis algorithm
Haario, H., Saksman, E., and Tamminen, J. (1999) · 1999
Earlier work this paper cites.
Rademacher processes and bounding the risk of function learning
Koltchinskii, V. and Panchenko, D. (2000) · 2000
Earlier work this paper cites.
Predictability, complexity, and learning
Bialek, W., Nemenman, I., and Tishby, N. (2001) · 2001
Earlier work this paper cites.
An adaptive metropolis algorithm
Haario, H., Saksman, E., and Tamminen, J. (2001) · 2001
Earlier work this paper cites.
Algebraic analysis for nonidentifiable learning machines
Watanabe, S. (2001) · 2001
Earlier work this paper cites.
Singularities in mixture models and upper bounds of stochastic complexity
Yamazaki, K. and Watanabe, S. (2003) · 2003
Earlier work this paper cites.
Resolution of singularities and the generalization error with Bayesian estimation for layered neural network
Aoyagi, M., Watanabe, S., et al. (2005) · 2005
Earlier work this paper cites.
Estimation of poles of zeta function in learning theory using Padé approximation
Iriguchi, R. and Watanabe, S. (2007) · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. (2009) · 2009
Earlier work this paper cites.
Algebraic Geometry and Statistical Learning Theory
Watanabe, S. (2009) · 2009
Earlier work this paper cites.
Asymptotic learning curve and renormalizable condition in statistical learning theory
Watanabe, S. (2010) · 2010
Earlier work this paper cites.
Bayesian learning via stochastic gradient Langevin dynamics
Welling, M. and Teh, Y. W. (2011) · 2011
Earlier work this paper cites.
The MNIST database of handwritten digit images for machine learning research
Deng, L. (2012) · 2012
Cited alongside, same era.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S. (2013) · 2013
Cited alongside, same era.
A Widely Applicable Bayesian Information Criterion
Watanabe, S. (2013) · 2013
Cited alongside, same era.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J. (2015) · 2015
Cited alongside, same era.
A guide to Bayesian model selection for ecologists
Hooten, M. B. and Hobbs, N. T. (2015) · 2015
Cited alongside, same era.
Deep residual learning for image recognition
What is the state of neural network pruning?
Blalock, D., Gonzalez Ortiz, J. J., Frankle, J., and Guttag, J. (2020) · 2020
Later among the works it cites.
Estimating the overdispersion in COVID-19 transmission using outbreak sizes outside China
Endo, A., Abbott, S., Kucharski, A. J., Funk, S., et al. (2020) · 2020
Later among the works it cites.
The Pile: an 800GB dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al. (2020) · 2020
Later among the works it cites.
Topological properties of the set of functions generated by neural networks of fixed size
Petersen, P. C., Raslan, M., and Voigtlaender, F. (2020) · 2020
Later among the works it cites.
Functional vs. parametric equivalence of ReLU networks
Phuong, M. and Lampert, C. H. (2020) · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Cited alongside, same era.
Sharp minima can generalize for deep nets
Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y. (2017) · 2017
Cited alongside, same era.
Exploring generalization in deep learning
Neyshabur, B., Bhojanapalli, S., Mcallester, D., and Srebro, N. (2017) · 2017
Cited alongside, same era.
Markov chain Monte Carlo methods for Bayesian data analysis in astronomy
Sharma, S. (2017) · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O. (2017) · 2017
Cited alongside, same era.
A convergence analysis of gradient descent for deep linear neural networks
Arora, S., Cohen, N., Golowich, N., and Hu, W. (2018) · 2018
Cited alongside, same era.
Gradient descent aligns the layers of deep linear networks
Ji, Z. and Telgarsky, M. (2018) · 2018
Cited alongside, same era.
A Bayesian neural network for toxicity prediction
Semenova, E., Williams, D. P., Afzal, A. M., and Lazic, S. E. (2020) · 2020
Later among the works it cites.
Scaling of sensory information in large neural populations shows signatures of information-limiting correlations
Kafashan, M., Jaffe, A. W., Chettih, S. N., Nogueira, R., Arandia-Romero, I., Harvey, C. D., Moreno-Bote, R., and Drugowitsch, J. (2021) · 2021
Later among the works it cites.
Shortformer: Better language modeling using shorter inputs
Press, O., Smith, N. A., and Lewis, M. (2021) · 2021
Later among the works it cites.
seaborn: statistical data visualization
Waskom, M. L. (2021) · 2021
Later among the works it cites.
Why neural networks find simple solutions: the many regularizers of geometric complexity
Dherin, B., Munn, M., Rosca, M., and Barrett, D. G. T. (2022) · 2022
Later among the works it cites.
TransformerLens
Nanda, N. and Bloom, J. (2022) · 2022
Later among the works it cites.
Deep Learning Is Singular, and That’s Good
Wei, S., Murfet, D., Gong, M., Li, H., Gell-Redman, J., and Quella, T. (2022) · 2022
Later among the works it cites.
Tensor programs V: Tuning large neural networks via zero-shot hyperparameter transfer
Yang, G., Hu, E. J., Babuschkin, I., Sidor, S., Liu, X., Farhi, D., Ryder, N., Pachocki, J., Chen, W., and Gao, J. (2022) · 2022
Later among the works it cites.
Sgd learning on neural networks: leap complexity and saddle-to-saddle dynamics
Abbe, E., Adserà, E. B., and Misiakiewicz, T. (2023) · 2023
Closest in time.
Incremental learning in diagonal linear networks
Berthier, R. (2023) · 2023
Closest in time.
Functional equivalence and path connectivity of reducible hyperbolic tangent networks
Farrugia-Roberts, M. (2023) · 2023
Closest in time.
You’re measuring model complexity wrong
Hoogland, J. and van Wingerden, S. (2023) · 2023
Closest in time.
Charting the topography of the neural network landscape with thermal-like noise
Jules, T., Brener, G., Kachman, T., Levi, N., and Bar-Sinai, Y. (2023) · 2023
Closest in time.
My criticism of singular learning theory
Skalse, J. (2023) · 2023
Closest in time.
Data selection for language models via importance resampling
Xie, S. M., Santurkar, S., Ma, T., and Liang, P. (2023) · 2023
Closest in time.
Consideration on the learning efficiency of multiple-layered neural networks with linear units
Aoyagi, M. (2024) · 2024
Closest in time.
The developmental landscape of in-context learning
Hoogland, J., Wang, G., Farrugia-Roberts, M., Carroll, L., Wei, S., and Murfet, D. (2024) · 2024
Closest in time.