Fetching the paper…
Reading the bibliography…
Solomonoff's general theory of inference and the Minimum Description Length principle formalize Occam's razor, and hold that a good model of data is a model that is good at losslessly compressing the data, including the cost of describing the model itself.
A mathematical theory of communication
C. Shannon · 1948
Earlier work this paper cites.
A formal theory of inductive inference
R. Solomonoff · 1964
Earlier work this paper cites.
Present Position and Potential Developments: Some Personal Views: Statistical Theory: The Prequential Approach
A. P. Dawid · 1984
Earlier work this paper cites.
Density estimation by stochastic complexity
J. Rissanen, T. Speed, and B. Yu · 1992
Earlier work this paper cites.
Keeping Neural Networks Simple by Minimizing the Description Length of the Weights
G. E. Hinton and D. Van Camp · 1993
Earlier work this paper cites.
The Risk Inflation Criterion for Multiple Regression
D. P. Foster and E. I. George · 1994
Earlier work this paper cites.
Discovering Neural Nets with Low Kolmogorov Complexity and High Generalization Capability
J. Schmidhuber · 1997
Earlier work this paper cites.
Gradient-Based Learning Applied to Document Recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Information-theoretic determination of minimax rates of convergence
A. Barron and Y. Yang · 1999
Earlier work this paper cites.
PAC-MDL Bounds
A. Blum and J. Langford · 2003
Earlier work this paper cites.
Information Theory, Inference, and Learning Algorithms
D. J. C. Mackay · 2003
Earlier work this paper cites.
Variational Learning and Bits-Back Coding: An Information-Theoretic View to Bayesian Learning
A. Honkela and H. Valpola · 2004
Earlier work this paper cites.
On the intelligibility of the universe and the notions of simplicity, complexity and irreducibility
G. J. Chaitin · 2007
Earlier work this paper cites.
The Minimum Description Length principle
P. D. Grünwald · 2007
Earlier work this paper cites.
On Universal Prediction and Bayesian Confirmation
M. Hutter · 2007
Earlier work this paper cites.
Information and complexity in statistical modeling
J. Rissanen · 2007
Cited alongside, same era.
An introduction to Kolmogorov complexity
M. Li and P. Vitányi · 2008
Cited alongside, same era.
Learning Multiple Layers of Features from Tiny Images
A. Krizhevsky · 2009
Cited alongside, same era.
Practical Variational Inference for Neural Networks
A. Graves · 2011
Cited alongside, same era.
Catching Up Faster by Switching Sooner: A predictive approach to adaptive estimation with an application to the AIC-BIC Dilemma
T. Van Erven, P. Grünwald, and S. De Rooij · 2012
Cited alongside, same era.
Auto-Encoding Variational Bayes
D. P. Kingma and M. Welling · 2013
Cited alongside, same era.
92.45% on CIFAR-10 in Torch, 2015
S. Zagoruyko · 2015
Later among the works it cites.
Degrees of Freedom in Deep Neural Networks
T. Gao and V. Jojic · 2016
Later among the works it cites.
Compression of Neural Machine Translation Models via Pruning
A. See, M.-T. Luong, and C. D. Manning · 2016
Later among the works it cites.
On the Emergence of Invariance and Disentangling in Deep Representations
A. Achille and S. Soatto · 2017
Later among the works it cites.
Computing Nonvacuous Generalization Bounds for Deep (Stochastic) Neural Networks with Many More Parameters than Training Data
G. K. Dziugaite and D. M. Roy · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Do Deep Nets Really Need to be Deep?
L. J. Ba and R. Caruana · 2014
Cited alongside, same era.
Auto-encoders: reconstruction versus compression
Y. Ollivier · 2014
Cited alongside, same era.
Very Deep Convolutional Networks for Large-Scale Image Recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Weight Uncertainty in Neural Networks
C. Blundell, J. Cornebise, K. Kavukcuoglu, and D. Wierstra · 2015
Cited alongside, same era.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Cited alongside, same era.
Automatic Differentiation Variational Inference
A. Kucukelbir, D. Tran, R. Ranganath, A. Gelman, and D. M. Blei · 2017
Later among the works it cites.
Bayesian compression for deep learning
C. Louizos, K. Ullrich, and M. Welling · 2017
Later among the works it cites.
Opening the Black Box of Deep Neural Networks via Information
R. Shwartz-Ziv and N. Tishby · 2017
Later among the works it cites.
Soft Weight-Sharing for Neural Network Compression
K. Ullrich, E. Meeds, and M. Welling · 2017
Later among the works it cites.
Margin-Aware Binarized Weight Networks for Image Classification
T.-B. Xu, P. Yang, X.-Y. Zhang, and C.-L. Liu · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2017
Later among the works it cites.
Stronger generalization bounds for deep nets via a compression approach
S. Arora, R. Ge, B. Neyshabur, and Y. Zhang · 2018
Closest in time.
Measuring the Intrinsic Dimension of Objective Landscapes
C. Li, H. Farkhoor, R. Liu, and J. Yosinski · 2018
Closest in time.
Pyvarinf : Variational Inference for PyTorch, 2018
C. Tallec and L. Blier · 2018
Closest in time.