Fetching the paper…
Reading the bibliography…
Memorization is worst-case generalization.
The Bell System Technical Journal
C. E. Shannon · 1948
Earlier work this paper cites.
The story of the binomial theorem
J. L. Coolidge · 1949
Earlier work this paper cites.
The perceptron: a probabilistic model for information storage and organization in the brain
F. Rosenblatt · 1958
Earlier work this paper cites.
Generalization and information storage in network of adaline’neurons’
B. Widrow · 1962
Earlier work this paper cites.
Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition
T. M. Cover · 1965
Earlier work this paper cites.
On the Uniform Convergence of Relative Frequencies of Events to Their Probabilities
V. N. Vapnik and A. Y. Chervonenkis · 1971
Earlier work this paper cites.
Dynamic connections in neural networks
J. A. Feldman · 1982
Earlier work this paper cites.
Learning Internal Representations by Error Propagation
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Occam’s razor
A. Blumer, A. Ehrenfeucht, D. Haussler, and M. K. Warmuth · 1987
Earlier work this paper cites.
Maximum storage capacity in neural networks
E. Gardner · 1987
Earlier work this paper cites.
The space of interactions in neural network models
E. Gardner · 1988
Earlier work this paper cites.
Learning Representations by Back-propagating Errors
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1988
Earlier work this paper cites.
Information theory, complexity and neural networks
Y. Abu-Mostafa · 1989
Earlier work this paper cites.
Training a 3-node neural network is np-complete
A. Blum and R. L. Rivest · 1989
Earlier work this paper cites.
Catastrophic Interference in Connectionist Networks: The Sequential Learning Problem
M. McCloskey and N. J. Cohen · 1989
Cited alongside, same era.
Connectionist models of recognition memory: constraints imposed by learning and forgetting functions
R. Ratcliff · 1990
Cited alongside, same era.
Measuring the VC-Dimension of a Learning Machine
V. N. Vapnik, E. Levin, and Y. L. Cun · 1994
Cited alongside, same era.
Neural networks: a systematic introduction
R. Rojas · 1996
Cited alongside, same era.
Neural Networks with Quadratic VC Dimension
P. Koiran and E. D. Sontag · 1997
Cited alongside, same era.
Vapnik-Chervonenkis dimension of recurrent neural networks
P. Koiran and E. D. Sontag · 1998
Cited alongside, same era.
Calculating the VC-dimension of decision trees
O. Asian, O. T. Yildiz, and E. Alpaydin · 2009
Later among the works it cites.
Rectified Linear Units Improve Restricted Boltzmann Machines
V. Nair and G. E. Hinton · 2010
Later among the works it cites.
Deep and Wide: Multiple Layers in Automatic Speech Recognition
N. Morgan · 2012
Later among the works it cites.
Understanding Machine Learning: From Theory to Algorithms
S. Shalev-Shwartz and S. Ben-David · 2014
Later among the works it cites.
Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Later among the works it cites.
Generalizing, Decoding, and Optimizing Support Vector Machine Classification
M. M. Krell · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Almost linear vc dimension bounds for piecewise polynomial networks
P. L. Bartlett, V. Maiorov, and R. Meir · 1999
Cited alongside, same era.
The nature of statistical learning theory
V. N. Vapnik · 2000
Cited alongside, same era.
Rademacher and Gaussian Complexities: Risk Bounds and Structural Results
P. L. Bartlett and S. Mendelson · 2001
Cited alongside, same era.
Information Theory, Inference, and Learning Algorithms
D. J. C. MacKay · 2003
Cited alongside, same era.
Online Passive-Aggressive Algorithms
K. Crammer, O. Dekel, J. Keshet, S. Shalev-Shwartz, and Y. Singer · 2006
Cited alongside, same era.
Ising models for networks of real neurons
G. Tkacik, E. Schneidman, I. Berry, J. Michael, and W. Bialek · 2006
Cited alongside, same era.
New one-class classifiers based on the origin separation approach
M. M. Krell and H. Wöhrle · 2015
Later among the works it cites.
Deep learning and the information bottleneck principle
N. Tishby and N. Zaslavsky · 2015
Later among the works it cites.
A Closer Look at Memorization in Deep Networks, jun 2017
D. Arpit, S. Jastrzȩbski, N. Ballas, D. Krueger, E. Bengio, M. S. Kanwal, T. Maharaj, A. Fischer, A. Courville, Y. Bengio, and S. Lacoste-Julien · 2017
Later among the works it cites.
Nearly-tight VC-dimension bounds for piecewise linear neural networks
N. Harvey, C. Liaw, and A. Mehrabian · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2017
Later among the works it cites.
A capacity scaling law for artificial neural networks
G. Friedland and M. Krell · 2018
Closest in time.
The helmholtz method: Using perceptual compression to reduce machine learning complexity
G. Friedland, J. Wang, R. Jia, and B. Li · 2018
Closest in time.