Fetching the paper…
Reading the bibliography…
In many applications, one works with neural network models trained by someone else.
Self-organized criticality: an explanation of
P. Bak, C. Tang, and K. Wiesenfeld · 1987
Earlier work this paper cites.
Convergent multiplicative processes repelled from zero: Power laws and truncated power laws
D. Sornette and R. Cont · 1997
Earlier work this paper cites.
Statistical Physics of Spin Glasses and Information Processing: An Introduction
H. Nishimori · 2001
Earlier work this paper cites.
Statistical mechanics of learning
A. Engel and C. P. L. Van den Broeck · 2001
Earlier work this paper cites.
Theory of Financial Risk and Derivative Pricing: From Statistical Physics to Risk Management
J. P. Bouchaud and M. Potters · 2003
Earlier work this paper cites.
Power laws, Pareto distributions and Zipf’s law
M. E. J. Newman · 2005
Earlier work this paper cites.
Critical phenomena in natural sciences: chaos, fractals, selforganization and disorder: concepts and tools
D. Sornette · 2006
Earlier work this paper cites.
Power-law distributions in empirical data
A. Clauset, C. R. Shalizi, and M. E. J. Newman · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Information capacity and transmission are maximized in balanced cortical networks with neuronal avalanches
W. L. Shew, H. Yang, S. Yu, R. Roy, and D. Plenz · 2011
Earlier work this paper cites.
Financial applications of random matrix theory: a short review
J. P. Bouchaud and M. Potters · 2011
Earlier work this paper cites.
Scale-invariant neuronal avalanche dynamics and the cut-off in size distributions
S. Yu, A. Klaus, H. Yang, and D. Plenz · 2014
Earlier work this paper cites.
powerlaw: A python package for analysis of heavy-tailed distributions
J. Alstott, E. Bullmore, and D. Plenz · 2014
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
O. Russakovsky et al · 2015
Cited alongside, same era.
Deep learning and the information bottleneck principle
N. Tishby and N. Zaslavsky · 2015
Cited alongside, same era.
Norm-based capacity control in neural networks
B. Neyshabur, R. Tomioka, and N. Srebro · 2015
Cited alongside, same era.
25 years of self-organized criticality: Concepts and controversies
N. W. Watkins, G. Pruessner, S. C. Chapman, N. B. Crosby, and H. J. Jensen · 2016
Cited alongside, same era.
Identity mappings in deep residual networks
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
On loss functions for deep neural networks in classification
https://pypi.org/project/WeightWatcher/
WeightWatcher, 2018 · 2018
Later among the works it cites.
A surprising linear relationship predicts test performance in deep networks
Q. Liao, B. Miranda, A. Banburski, J. Hidary, and T. Poggio · 2018
Later among the works it cites.
The singular values of convolutional layers
H. Sedghi, V. Gupta, and P. M. Long · 2018
Later among the works it cites.
Traditional and heavy-tailed self regularization in neural network models
C. H. Martin and M. W. Mahoney · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke et al · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Janocha and W. M. Czarnecki · 2017
Cited alongside, same era.
Cleaning large correlation matrices: tools from random matrix theory
J. Bun, J.-P. Bouchaud, and M. Potters · 2017
Cited alongside, same era.
Opening the black box of deep neural networks via information
R. Shwartz-Ziv and N. Tishby · 2017
Cited alongside, same era.
A survey of model compression and acceleration for deep neural networks
Y. Cheng, D. Wang, P. Zhou, and T. Zhang · 2017
Cited alongside, same era.
A. Vaswani et al · 2017
Cited alongside, same era.
C. H. Martin and M. W. Mahoney · 2017
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
P. Bartlett, D. J. Foster, and M. Telgarsky · 2017
Cited alongside, same era.
T. Wolf et al · 2019
Later among the works it cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
M. Belkin, D. Hsu, S. Ma, and S. Mandal · 2019
Later among the works it cites.
Statistical mechanics methods for discovering knowledge from modern production quality neural networks
C. H. Martin and M. W. Mahoney · 2019
Later among the works it cites.
Heavy-tailed Universality predicts trends in test accuracies for very large pre-trained deep neural networks
C. H. Martin and M. W. Mahoney · 2020
Closest in time.
Multiplicative noise and heavy tails in stochastic optimization
L. Hodgkinson and M. W. Mahoney · 2020
Closest in time.
Statistical mechanics of deep learning
Y. Bahri, J. Kadmon, J. Pennington, S. Schoenholz, J. Sohl-Dickstein, and S. Ganguli · 2020
Closest in time.
Classifying the classifier: dissecting the weight space of neural networks
G. Eilertsen, D. Jönsson, T. Ropinski, J. Unger, and A. Ynnerman · 2020
Closest in time.
Predicting neural network accuracy from weights
T. Unterthiner, D. Keysers, S. Gelly, O. Bousquet, and I. Tolstikhin · 2020
Closest in time.