Fetching the paper…
Reading the bibliography…
We describe an approach to understand the peculiar and counterintuitive generalization properties of deep neural networks.
Principles of Neurodynamics
F. Rosenblatt · 1962
Earlier work this paper cites.
Statistical Mechanics of Neural Networks
J. D. Cowan · 1967
Earlier work this paper cites.
The existence of persistent states in the brain
W. A. Little · 1974
Earlier work this paper cites.
Solutions of Ill-Posed Problems
A.N. Tikhonov and V.Y. Arsenin · 1977
Earlier work this paper cites.
The random energy model
B. Derrida · 1980
Earlier work this paper cites.
Neural networks and physical systems with emergent collective computational abilities
J. J. Hopfield · 1982
Earlier work this paper cites.
A theory of the learnable
L. G. Valiant · 1984
Earlier work this paper cites.
Learning and relearning in Boltzmann machines
G. E. Hinton and T. J. Sejnowski · 1986
Earlier work this paper cites.
Exhaustive thermodynamical analysis of Boolean learning networks
P. Carnevali and S. Patarnello · 1987
Earlier work this paper cites.
The truncated SVD as a method for regularization
P. C. Hansen · 1987
Earlier work this paper cites.
Learning representations by back-propagating errors
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1988
Earlier work this paper cites.
Statistical mechanics of neural networks
H. Sompolinsky · 1988
Earlier work this paper cites.
Three unfinished works on the optimal storage capacity of networks
E. Gardner and B. Derrida · 1989
Earlier work this paper cites.
A statistical approach to learning and generalization in layered neural networks
E. Levin, N. Tishby, and S. A. Solla · 1990
Earlier work this paper cites.
Statistical mechanics of a multilayered neural network
E. Barkai, D. Hansel, and I. Kanter · 1990
Earlier work this paper cites.
First-order transition to perfect generalization in a neural network with binary synapses
G. Györgyi · 1990
Earlier work this paper cites.
Statistical mechanics and phase transitions in clustering
K. Rose, E. Gurewitz, and G. C. Fox · 1990
Earlier work this paper cites.
Learning from examples in large neural networks
H. Sompolinsky, N. Tishby, and H. S. Seung · 1990
Earlier work this paper cites.
Constrained neural networks for pattern recognition
S. A. Solla and Y. LeCun · 1991
Earlier work this paper cites.
How tight are the Vapnik-Chervonenkis bounds?
D. Cohn and G. Tesauro · 1992
Earlier work this paper cites.
Statistical mechanics of learning from examples
H. S. Seung, H. Sompolinsky, , and N. Tishby · 1992
Earlier work this paper cites.
Four types of learning curves
S.-I. Amari, N. Fujita, and S. Shinomoto · 1992
Earlier work this paper cites.
Generalization in a two-layer neural network
H. Schwarze, M. Opper, and W. Kinzel · 1992
Earlier work this paper cites.
Memorization without generalization in a multilayered neural network
D. Hansel, G. Mato, and C. Meunier · 1992
Earlier work this paper cites.
Learning in linear neural networks: The validity of the annealed approximation
S. A. Solla and E. Levin · 1992
Earlier work this paper cites.
The statistical mechanics of learning a rule
T. L. H. Watkin, A. Rau, and M. Biehl · 1993
Earlier work this paper cites.
Statistical mechanics of learning in a large committee machine
H. Schwarze and J. Hertz · 1993
Earlier work this paper cites.
Learning a rule in a multilayer neural network
H. Schwarze · 1993
Earlier work this paper cites.
Statistical mechanics calculation of Vapnik-Chervonenkis bounds for perceptrons
A. Engel and W. Fink · 1993
Earlier work this paper cites.
Measuring the VC-dimension of a learning machine
V. Vapnik, E. Levin, and Y. Le Cun · 1994
Earlier work this paper cites.
Learning and generalization in a two-layer neural network: the role of the Vapnik-Chervonenkis dimension
M. Opper · 1994
Earlier work this paper cites.
Reliability of replica symmetry for the generalization problem of a toy multilayer neural network
A. Engel and L. Reimers · 1994
Earlier work this paper cites.
Perfect loss of generalization due to noise in k = 2 k=2 parity machines
Y. Kabashima · 1994
Earlier work this paper cites.
Neural Networks for Pattern Recognition
C. M. Bishop · 1995
Cited alongside, same era.
On the consequences of the statistical mechanics theory of learning curves for the model selection problem
M. J. Kearns · 1995
Cited alongside, same era.
Storage capacity and generalization for the reversed-wedge Ising perceptron
G. J. Bex, R. Serneels, and C. van den Broeck · 1995
Cited alongside, same era.
On-line backpropagation in two-layered neural networks
P. Riegler and M. Biehl · 1995
Cited alongside, same era.
Perceptron learning: The largest version space
M. Biehl and M. Opper · 1995
Cited alongside, same era.
On-line learning in the committee machine
M. Copelli and N. Caticha · 1995
Cited alongside, same era.
Statistical mechanics of complex neural systems and high dimensional data
M. Advani, S. Lahiri, and S. Ganguli · 2013
Later among the works it cites.
High-dimensional random fields and random matrix theory
Y. V Fyodorov · 2013
Later among the works it cites.
Training convolutional networks with noisy labels
S. Sukhbaatar, J. Bruna, M. Paluri, L. Bourdev, and R. Fergus · 2014
Later among the works it cites.
Qualitatively characterizing neural network optimization problems
I. J. Goodfellow, O. Vinyals, and A. M. Saxe · 2014
Later among the works it cites.
Explorations on high dimensional landscapes
L. Sagun, V. U. Guney, G. Ben Arous, and Y. LeCun · 2014
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rigorous learning curve bounds from statistical mechanics
D. Haussler, M. Kearns, H. S. Seung, and N. Tishby · 1996
Cited alongside, same era.
A numerical study on learning curves in stochastic multilayer feedforward networks
K.-R. Müller, M. Finke, N. Murata, K. Schulten, and S.-I. Amari · 1996
Cited alongside, same era.
Transient dynamics of on-line learning in two-layered neural networks
M. Biehl, P. Riegler, and C. Wohler · 1996
Cited alongside, same era.
On-line learning in parity machines
R. Simonetti and N. Caticha · 1996
Cited alongside, same era.
Equivalence between learning in noisy perceptrons and tree committee machines
M. Copelli, O. Kinouchi, and N. Caticha · 1996
Cited alongside, same era.
For valid generalization, the size of the weights is more important than the size of the network
P. L. Bartlett · 1997
Cited alongside, same era.
Later among the works it cites.
The loss surfaces of multilayer networks
A. Choromanska, M. Henaff, M. Mathieu, G. Ben Arous, and Y. LeCun · 2014
Later among the works it cites.
On the saddle point problem for non-convex optimization
R. Pascanu, Y. N. Dauphin, S. Ganguli, and Y. Bengio · 2014
Later among the works it cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Y. N. Dauphin, R. Pascanu, C. Gulcehre, K. Cho, S. Ganguli, and Y. Bengio · 2014
Later among the works it cites.
In search of the real inductive bias: on the role of implicit regularization in deep learning
B. Neyshabur, R. Tomioka, and N. Srebro · 2014
Later among the works it cites.
The unreasonable effectiveness of noisy data for fine-grained recognition
J. Krause, B. Sapp, A. Howard, H. Zhou, A. Toshev, T. Duerig, J. Philbin, and L. Fei-Fei · 2015
Later among the works it cites.
Making Vapnik-Chervonenkis bounds accurate
L. Bottou · 2015
Later among the works it cites.
The effect of gradient noise on the energy landscape of deep networks
P. Chaudhari and S. Soatto · 2015
Later among the works it cites.
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Later among the works it cites.
Universal adversarial perturbations
S.-M. Moosavi-Dezfooli, A. Fawzi, O. Fawzi, and P. Frossard · 2016
Later among the works it cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2016
Later among the works it cites.
On large-batch training for deep learning: generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2016
Later among the works it cites.
Statistical mechanics of optimal convex inference in high dimensions
M. Advani and S. Ganguli · 2016
Later among the works it cites.
Statistical physics of inference: thresholds and algorithms
L. Zdeborová and F. Krzakala · 2016
Later among the works it cites.
Deep learning is robust to massive label noise
D. Rolnick, A. Veit, S. Belongie, and N. Shavit · 2017
Closest in time.
Generalization in deep learning
K. Kawaguchi, L. P. Kaelbling, and Y. Bengio · 2017
Closest in time.
Sharp minima can generalize for deep nets
L. Dinh, R. Pascanu, S. Bengio, and Y. Bengio · 2017
Closest in time.
C. H. Martin and M. W. Mahoney · 2017
Closest in time.
A capacity scaling law for artificial neural networks
G. Friedland and M. Krell · 2017
Closest in time.
Spectrally-normalized margin bounds for neural networks
P. Bartlett, D. J. Foster, and M. Telgarsky · 2017
Closest in time.
Opening the black box of deep neural networks via information
R. Shwartz-Ziv and N. Tishby · 2017
Closest in time.
A surprising linear relationship predicts test performance in deep networks
Q. Liao, B. Miranda, A. Banburski, J. Hidary, and T. Poggio · 2018
Closest in time.
Theory IIIb: Generalization in deep networks
T. Poggio, Q. Liao, B. Miranda, A. Banburski, X. Boix, and J. Hidary · 2018
Closest in time.
On the computational inefficiency of large batch sizes for stochastic gradient descent
N. Golmant, N. Vemuri, Z. Yao, V. Feinberg, A. Gholami, K. Rothauge, M. W. Mahoney, and J. Gonzalez · 2018
Closest in time.
Measuring the effects of data parallelism on neural network training
C. J. Shallue, J. Lee, J. Antognini, J. Sohl-Dickstein, R. Frostig, and G. E. Dahl · 2018
Closest in time.
C. H. Martin and M. W. Mahoney · 2018
Closest in time.
Traditional and heavy-tailed self regularization in neural network models
C. H. Martin and M. W. Mahoney · 2019
Closest in time.
C. H. Martin and M. W. Mahoney · 2019
Closest in time.