Fetching the paper…
Reading the bibliography…
Random Matrix Theory (RMT) is applied to analyze the weight matrices of Deep Neural Networks (DNNs), including both production quality, pre-trained models such as AlexNet and Inception, and smaller models trained from scratch, such as LeNet5 and a miniature-AlexNet.
Efficient backprop in neural networks: Tricks of the trade
Y. LeCun, L. Bottou, and G. Orr · 1988
Earlier work this paper cites.
Statistical mechanics of learning from examples
H. S. Seung, H. Sompolinsky, , and N. Tishby · 1992
Earlier work this paper cites.
The statistical mechanics of learning a rule
T. L. H. Watkin, A. Rau, and M. Biehl · 1993
Earlier work this paper cites.
Measuring the VC-dimension of a learning machine
V. Vapnik, E. Levin, and Y. Le Cun · 1994
Earlier work this paper cites.
Rigorous learning curve bounds from statistical mechanics
D. Haussler, M. Kearns, H. S. Seung, and N. Tishby · 1996
Earlier work this paper cites.
Universality classes for extreme-value statistics
J.-P. Bouchaud and M. Mézard · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Rational decisions, random matrices and spin glasses
S. Galluccio, J.-P. Bouchaud, and M. Potters · 1998
Earlier work this paper cites.
Noise dressing of financial correlation matrices
L. Laloux, P. Cizeau, J.-P. Bouchaud, and M. Potters · 1999
Earlier work this paper cites.
Pattern Classification
R. O. Duda, P. E. Hart, and D. G. Stork · 2001
Earlier work this paper cites.
Statistical mechanics of learning
A. Engel and C. P. L. Van den Broeck · 2001
Earlier work this paper cites.
On the distribution of the largest eigenvalue in principal components analysis
I. M. Johnstone · 2001
Earlier work this paper cites.
Collective origin of the coexistence of apparent RMT noise and factors in large sample correlation matrices
Y. Malevergne and D. Sornette · 2002
Earlier work this paper cites.
Random matrix theory and wireless communications
A. M. Tulino and S. Verdú · 2004
Earlier work this paper cites.
Problems with fitting to the power-law distribution
M. L. Goldstein, S. A. Morris, and G. G. Yen · 2004
Earlier work this paper cites.
Random matrix theory
A. Edelman and N. R. Rao · 2005
Earlier work this paper cites.
Recent results about the largest eigenvalue of random covariance matrices and statistical applications
N. El Karoui · 2005
Earlier work this paper cites.
Random matrix theory and financial correlations
L. Laloux, P. Cizeau, M. Potters, and J.-P. Bouchaud · 2005
Cited alongside, same era.
Empirical distributions of stock returns: between the stretched exponential and the power law?
Y. Malevergne, V. Pisarenko, and D. Sornette · 2005
Cited alongside, same era.
Critical phenomena in natural sciences: chaos, fractals, selforganization and disorder: concepts and tools
D. Sornette · 2006
Cited alongside, same era.
On the top eigenvalue of heavy-tailed random matrices
G. Biroli, J.-P. Bouchaud, and M. Potters · 2007
Cited alongside, same era.
Extreme value problems in random matrix theory and other disordered systems
G. Biroli, J.-P. Bouchaud, and M. Potters · 2007
Cited alongside, same era.
Parameter estimation for power-law distributions by maximum likelihood methods
H. Bauke · 2007
Random matrix theory in statistics: a review
D. Paul and A. Aue · 2014
Later among the works it cites.
Limit theory for the largest eigenvalues of sample covariance matrices with heavy-tails
R. A. Davis, O. Pfaffel, and R. Stelzer · 2014
Later among the works it cites.
powerlaw: A python package for analysis of heavy-tailed distributions
J. Alstott, E. Bullmore, and D. Plenz · 2014
Later among the works it cites.
Power-law distributions in binned empirical data
Y. Virkar and A. Clauset · 2014
Later among the works it cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2016
Later among the works it cites.
On large-batch training for deep learning: generalization gap and sharp minima
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The distributions of random matrix theory and their applications
C. A. Tracy and H. Widom · 2009
Cited alongside, same era.
Poisson convergence for the largest eigenvalues of heavy tailed random matrices
A. Auffinger, G. Ben Arous, and S. Péché · 2009
Cited alongside, same era.
Z. Burda and J. Jurkiewicz · 2009
Cited alongside, same era.
Power-law distributions in empirical data
A. Clauset, C. R. Shalizi, and M. E. J. Newman · 2009
Cited alongside, same era.
Financial applications of random matrix theory: a short review
J. P. Bouchaud and M. Potters · 2011
Cited alongside, same era.
Statistical analyses support power law distributions found in neuronal avalanches
A. Klaus, S. Yu, and D. Plenz · 2011
Cited alongside, same era.
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2016
Later among the works it cites.
Beyond universality in random matrix theory
A. Edelman, A. Guionnet, and S. Péché · 2016
Later among the works it cites.
Extreme eigenvalues of sparse, heavy tailed random matrices
A. Auffinger and S. Tang · 2016
Later among the works it cites.
C. H. Martin and M. W. Mahoney · 2017
Later among the works it cites.
Regularization for deep learning: A taxonomy
J. Kukačka, V. Golkov, and D. Cremers · 2017
Later among the works it cites.
E. Hoffer, I. Hubara, and D. Soudry · 2017
Later among the works it cites.
Accurate, large minibatch SGD: Training ImageNet in 1 hour
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Later among the works it cites.
Cleaning large correlation matrices: tools from random matrix theory
J. Bun, J.-P. Bouchaud, and M. Potters · 2017
Later among the works it cites.
Fitting power-laws in empirical data with estimators that work for all exponents
R. Hanel, B. Corominas-Murtra, B. Liu, and S. Thurner · 2017
Later among the works it cites.
C. H. Martin and M. W. Mahoney · 2018
Later among the works it cites.
Heavy-tailed Universality predicts trends in test accuracies for very large pre-trained deep neural networks
C. H. Martin and M. W. Mahoney · 2019
Closest in time.