Fetching the paper…
Reading the bibliography…
To understand better good generalization performance in state-of-the-art neural network (NN) models, and in particular the success of the ALPHAHAT metric based on Heavy-Tailed Self-Regularization (HT-SR) theory, we analyze of a corpus of models that was made publicly-available for a contest to predict the generalization accuracy of NNs.
The interpretation of interaction in contingency tables
E. H. Simpson · 1951
Earlier work this paper cites.
Sex bias in graduate admissions: Data from Berkeley
P. J. Bickel, E. A. Hammel, and J. W. O’Connell · 1975
Earlier work this paper cites.
For valid generalization, the size of the weights is more important than the size of the network
P. L. Bartlett · 1997
Earlier work this paper cites.
Statistical mechanics of learning
A. Engel and C. P. L. Van den Broeck · 2001
Earlier work this paper cites.
Theory of Financial Risk and Derivative Pricing: From Statistical Physics to Risk Management
J. P. Bouchaud and M. Potters · 2003
Earlier work this paper cites.
Power laws, Pareto distributions and Zipf’s law
M. E. J. Newman · 2005
Earlier work this paper cites.
Analysis of power laws, shape collapses, and neural complexity: New techniques and MATLAB support via the NCC toolbox
N. Marshall, N. M. Timme, N. Bennett, M. Ripp, E. Lautzenhiser, and J. M. Beggs · 2005
Earlier work this paper cites.
Probability distributions in complex systems
D. Sornette · 2007
Earlier work this paper cites.
Power-law distributions in empirical data
A. Clauset, C. R. Shalizi, and M. E. J. Newman · 2009
Earlier work this paper cites.
Ecological correlations and the behavior of individuals
W. S. Robinson · 2009
Earlier work this paper cites.
Causality: Models, Reasoning and Inference
J. Pearl · 2009
Earlier work this paper cites.
Being critical of criticality in the brain
J. M. Beggs and N. Timme · 2012
Cited alongside, same era.
Simpson’s paradox in psychological science: a practical guide
R. A. Kievit, W. E. Frankenhuis, L. J. Waldorp, and D. Borsboom · 2013
Cited alongside, same era.
powerlaw: A python package for analysis of heavy-tailed distributions
J. Alstott, E. Bullmore, and D. Plenz · 2014
Cited alongside, same era.
Norm-based capacity control in neural networks
B. Neyshabur, R. Tomioka, and N. Srebro · 2015
Cited alongside, same era.
Statistical physics of inference: thresholds and algorithms
L. Zdeborová and F. Krzakala · 2016
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
P. Bartlett, D. J. Foster, and M. Telgarsky · 2017
Fantastic generalization measures and where to find them
Y. Jiang, B. Neyshabur, H. Mobahi, D. Krishnan, and S. Bengio · 2019
Later among the works it cites.
On the interplay between noise and curvature and its effect on optimization and generalization
V. Thomas, F. Pedregosa, B. van Merrienboer, P.-A. Mangazol, Y. Bengio, and N. Le Roux · 2019
Later among the works it cites.
Heavy-tailed Universality predicts trends in test accuracies for very large pre-trained deep neural networks
C. H. Martin and M. W. Mahoney · 2020
Later among the works it cites.
NeurIPS 2020 competition: Predicting generalization in deep learning (version 1.0)
Y. Jiang, P. Foret, S. Yak, D. M. Roy, H. Mobahi, G. K. Dziugaite, S. Bengio, S. Gunasekar, I. Guyon, and B. Neyshabur · 2020
Later among the works it cites.
NeurIPS 2020 competition: Predicting generalization in deep learning (version 1.1)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
https://pypi.org/project/WeightWatcher/
WeightWatcher, 2018 · 2018
Cited alongside, same era.
Stronger generalization bounds for deep nets via a compression approach
S. Arora, R. Ge, B. Neyshabur, and Y. Zhang · 2018
Cited alongside, same era.
Approximate Fisher information matrix to characterise the training of deep neural networks
Z. Liao, T. Drummond, I. Reid, and G. Carneiro · 2018
Cited alongside, same era.
Traditional and heavy-tailed self regularization in neural network models
C. H. Martin and M. W. Mahoney · 2019
Cited alongside, same era.
Y. Jiang, P. Foret, S. Yak, D. M. Roy, H. Mobahi, G. K. Dziugaite, S. Bengio, S. Gunasekar, I. Guyon, and B. Neyshabur · 2020
Later among the works it cites.
Statistical mechanics of deep learning
Y. Bahri, J. Kadmon, J. Pennington, S. Schoenholz, J. Sohl-Dickstein, and S. Ganguli · 2020
Later among the works it cites.
In search of robust measures of generalization
G. K. Dziugaite, A. Drouin, B. Neal, N. Rajkumar, E. Caballero, L. Wang, I. Mitliagkas, and D. M. Roy · 2020
Later among the works it cites.
Predicting trends in the quality of state-of-the-art neural networks without access to training or testing data
C. H. Martin, T. S. Peng, and M. W. Mahoney · 2021
Closest in time.
Implicit self-regularization in deep neural networks: Evidence from random matrix theory and implications for learning
C. H. Martin and M. W. Mahoney · 2021
Closest in time.
Y. Yang, R. Theisen, L. Hodgkinson, J. E. Gonzalez, K. Ramchandran, C. H. Martin, and M. W. Mahoney · 2022
Closest in time.