On the complexity of finite sequences
Abraham Lempel and Jacob Ziv · 1976
Earlier work this paper cites.
Modeling by shortest data description
Jorma Rissanen · 1978
Earlier work this paper cites.
Occam’s razor
Anselm Blumer, Andrzej Ehrenfeucht, David Haussler, and Manfred K Warmuth · 1987
Earlier work this paper cites.
What size net gives valid generalization?
Eric B Baum and David Haussler · 1989
Earlier work this paper cites.
Generalization and parameter estimation in feedforward nets: Some experiments
Nelson Morgan and Hervé Bourlard · 1990
Earlier work this paper cites.
A simple weight decay can improve generalization
Anders Krogh and John A Hertz · 1992
Earlier work this paper cites.
Keeping the neural networks simple by minimizing the description length of the weights
Geoffrey E. Hinton and Drew van Camp · 1993
Earlier work this paper cites.
The relationship between pac, the statistical physics framework, the bayesian framework, and the vc framework
David H Wolpert and R Waters · 1994
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Discovering neural nets with low kolmogorov complexity and high generalization capability
Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Boolean functions with low average sensitivity depend on few coordinates
Ehud Friedgut · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Some pac-bayesian theorems
David A McAllester · 1998
Earlier work this paper cites.
Bounds for averaging classifiers
John Langford and Matthias Seeger · 2001
Earlier work this paper cites.
On a generalization complexity measure for boolean functions
Leonardo Franco and Martin Anthony · 2004
Earlier work this paper cites.
Gaussian processes in machine learning
Carl Edward Rasmussen · 2004
Earlier work this paper cites.
Generalization ability of boolean functions implemented in feedforward neural networks
Leonardo Franco · 2006
Earlier work this paper cites.
On early stopping in gradient descent learning
Yuan Yao, Lorenzo Rosasco, and Andrea Caponnetto · 2007
Earlier work this paper cites.
Kernel methods for deep learning
Youngmin Cho and Lawrence K Saul · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Elements of information theory
Thomas M Cover and Joy A Thomas · 2012
Earlier work this paper cites.
GPy: A gaussian process framework in python
GPy · 2012
Earlier work this paper cites.
Robustness and generalization
Huan Xu and Shie Mannor · 2012
Earlier work this paper cites.
On the non-randomness of maximum lempel ziv complexity sequences of finite size
E Estevez-Rams, R Lora Serrano, B Aragón Fernández, and I Brito Reyes · 2013
Earlier work this paper cites.
No free lunch versus Occam’s razor in supervised learning
Tor Lattimore and Marcus Hutter · 2013
Earlier work this paper cites.