Fetching the paper…
Reading the bibliography…
This is an up-to-date introduction to and overview of the Minimum Description Length (MDL) Principle, a theory of inductive inference that can be applied to general problems in statistics, machine learning and pattern recognition.
Safe testing, 2019
P. Grünwald, R. de Heide, and W. Koolen · 1906
Earlier work this paper cites.
An information measure for classification
C.S. Wallace and D.M. Boulton · 1968
Earlier work this paper cites.
Modeling by the shortest data description
J. Rissanen · 1978
Earlier work this paper cites.
Present position and potential developments: Some personal views, statistical theory, the prequential approach
A.P. Dawid · 1984
Earlier work this paper cites.
Universal coding, information, prediction and estimation
J. Rissanen · 1984
Earlier work this paper cites.
Universal sequential coding of single messages
Yu. M. Shtarkov · 1987
Earlier work this paper cites.
Stochastic Complexity in Statistical Inquiry
J. Rissanen · 1989
Earlier work this paper cites.
Minimum complexity density estimation
A.R. Barron and T.M. Cover · 1991
Earlier work this paper cites.
Elements of Information Theory
T.M. Cover and J.A. Thomas · 1991
Earlier work this paper cites.
Complexity of models
J. Rissanen · 1991
Earlier work this paper cites.
Keeping the neural networks simple by minimizing the description length of the weights
Geoffrey E Hinton and Drew Van Camp · 1993
Earlier work this paper cites.
Learning Bayesian belief networks: an approach based on the MDL principle
Wai Lam and Fahiem Bacchus · 1994
Earlier work this paper cites.
A characterization of the Dirichlet distribution with application to learning Bayesian networks
Dan Geiger and David Heckerman · 1995
Earlier work this paper cites.
Learning Bayesian networks: The combination of knowledge and statistical data
David Heckerman, Dan Geiger, and David Chickering · 1995
Earlier work this paper cites.
Bayes factors
Robert E Kass and Adrian E Raftery · 1995
Earlier work this paper cites.
Fisher information and stochastic complexity
J. Rissanen · 1996
Earlier work this paper cites.
Flat minima
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
The minimum description length principle in coding and modeling
A. Barron, J. Rissanen, and B. Yu · 1998
Earlier work this paper cites.
Bayes factors and marginal distributions in invariant situations
James O Berger, Luis R Pericchi, and Julia A Varshavsky · 1998
Earlier work this paper cites.
Robustly minimax codes for universal data compression
J. Takeuchi and A. R. Barron · 1998
Earlier work this paper cites.
Viewing all models as “probabilistic”
P. D. Grünwald · 1999
Earlier work this paper cites.
Estimation of Mixture Models
J.Q. Li · 1999
Earlier work this paper cites.
Mixture density estimation
J.Q. Li and A.R. Barron · 2000
Earlier work this paper cites.
Counting probability distributions: Differential geometry and model selection
I.J. Myung, V. Balasubramanian, and M.A. Pitt · 2000
Earlier work this paper cites.
Occam’s razor
C.E. Rasmussen and Z. Ghahramani · 2000
Earlier work this paper cites.
The last-step minimax algorithm
Eiji Takimoto and Manfred Warmuth · 2000
Earlier work this paper cites.
Minimum description length induction, Bayesianism, and Kolmogorov complexity
P.M.B. Vitányi and M. Li · 2000
Earlier work this paper cites.
The Elements of Statistical Learning: Data Mining, Inference and Prediction
T. Hastie, R. Tibshirani, and J. Friedman · 2001
Earlier work this paper cites.
(Not) bounding the true error
John Langford and Rich Caruana · 2002
Earlier work this paper cites.
PAC-MDL bounds
A. Blum and J. Langford · 2003
Earlier work this paper cites.
Unified conditional frequentist and Bayesian testing of composite hypotheses
Sarat C. Dass and James O. Berger · 2003
Earlier work this paper cites.
PAC-Bayesian stochastic model selection
D. McAllester · 2003
Earlier work this paper cites.
Probabilistic network construction using the minimum description length principle
Remco R. Bouckaert · 2005
Cited alongside, same era.
Advances in minimum description length: Theory and applications
Peter D Grünwald, In Jae Myung, and Mark A Pitt · 2005
Cited alongside, same era.
Can the strengths of AIC and BIC be shared? A conflict between model indentification and regression estimation
Y. Yang · 2005
Cited alongside, same era.
Prediction, Learning and Games
N. Cesa-Bianchi and G. Lugosi · 2006
Cited alongside, same era.
The geometry of proper scoring rules
A Philip Dawid · 2007
Cited alongside, same era.
Catching up faster in Bayesian model selection and model averaging
T. van Erven, P.D. Grünwald, and S. de Rooij · 2007
Cited alongside, same era.
Information theoretic validity of penalized likelihood
Sabyasachi Chatterjee and Andrew Barron · 2014
Later among the works it cites.
Robust learning of inhomogeneous PMMs
Ralf Eggeling, Teemu Roos, Petri Myllymäki, and Ivo Grosse · 2014
Later among the works it cites.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Later among the works it cites.
Inferring intra-motif dependencies of DNA binding sites from chip-seq data
Ralf Eggeling, Teemu Roos, Petri Myllymäki, and Ivo Grosse · 2015
Later among the works it cites.
Summarizing and understanding large graphs
Danai Koutra, U Kang, Jilles Vreeken, and Christos Faloutsos · 2015
Later among the works it cites.
Achievability of asymptotic minimax regret by horizon-dependent and horizon-independent strategies
Kazuho Watanabe and Teemu Roos · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The Minimum Description Length Principle
P. Grünwald · 2007
Cited alongside, same era.
Suboptimal behavior of Bayes and MDL in classification under misspecification
P. Grünwald and J. Langford · 2007
Cited alongside, same era.
A linear-time algorithm for computing the multinomial stochastic complexity
Petri Kontkanen and Petri Myllymäki · 2007
Cited alongside, same era.
Information and Complexity in Statistical Modeling
J. Rissanen · 2007
Cited alongside, same era.
The MDL principle, penalized likelihoods, and statistical risk
A Barron, Cong Huang, J Li, and Xi Luo · 2008
Cited alongside, same era.
MDL procedures with L1 penalty and their statistical risk
Andrew Barron and Xi Luo · 2008
Cited alongside, same era.
Later among the works it cites.
The Safe–Bayesian Lasso
R de Heide · 2016
Later among the works it cites.
Barron and Cover’s theory in supervised learning and its application to Lasso
Masanori Kawakita and Jun-ichi Takeuchi · 2016
Later among the works it cites.
Relations between the conditional normalized maximum likelihood distributions and the latent information priors
Mutsuki Kojima and Fumiyasu Komaki · 2016
Later among the works it cites.
Harold Jeffreys’ default Bayes factor hypothesis tests: Explanation, extension, and application in psychology
A. Ly, J. Verhagen, and E.J. Wagenmakers · 2016
Later among the works it cites.
Robust sequential prediction in linear regression with Student’s t t -distribution
Jussi Määttä and Teemu Roos · 2016
Later among the works it cites.
Subset selection in linear regression using sequentially normalized least squares: Asymptotic theory
Jussi Määttä, Daniel F. Schmidt, and Teemu Roos · 2016
Later among the works it cites.
Structure selection for convolutive non-negative matrix factorization using normalized maximum likelihood coding
Atsushi Suzuki, Kohei Miyaguchi, and Kenji Yamanishi · 2016
Later among the works it cites.
Jeffreys’ and BDeu priors for model selection
Joe Suzuki · 2016
Later among the works it cites.
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Later among the works it cites.
Inconsistency of Bayesian inference for misspecified linear models, and a proposal for repairing it
P. Grünwald and T. van Ommen · 2017
Later among the works it cites.
An upper bound on normalized maximum likelihood codes for gaussian mixture models
So Hirai and Kenji Yamanishi · 2017
Later among the works it cites.
Sparse graphical modeling via stochastic complexity
Kohei Miyaguchi, Shin Matsushima, and Kenji Yamanishi · 2017
Later among the works it cites.
Decomposed normalized maximum likelihood codelength criterion for selecting hierarchical latent variable models
Tianyi Wu, Shinya Sugawara, and Kenji Yamanishi · 2017
Later among the works it cites.
On model selection, Bayesian networks, and the Fisher information integral
Yuan Zou and Teemu Roos · 2017
Later among the works it cites.
Finite-sample risk bounds for maximum likelihood estimation with arbitrary penalties
W. Brinda and J. Klusowski · 2018
Later among the works it cites.
Causal inference by compression
K. Budhathoki, J. Vreeken, and J. Origo · 2018
Later among the works it cites.
Safe probability
Peter Grünwald · 2018
Later among the works it cites.
High-dimensional penalty selection via minimum description length principle
K. Miyaguchi and K. Yamanishi · 2018
Later among the works it cites.
Almost the best of three worlds: Risk, consistency and optional stopping for the switch criterion in nested model selection
S. van der Pas and P D Grünwald · 2018
Later among the works it cites.
Quotient normalized maximum likelihood criterion for learning Bayesian network structures
Tomi Silander, Janne Leppä-aho, Elias Jääsaari, and Teemu Roos · 2018
Later among the works it cites.
Universal prediction: a philosophical investigation
Tom Florian Sterkenburg · 2018
Later among the works it cites.
Exact calculation of normalized maximum likelihood code length using Fourier analysis
Atsushi Suzuki and Kenji Yamanishi · 2018
Later among the works it cites.
Compressibility and generalization in large-scale deep learning
Wenda Zhou, Victor Veitch, Morgane Austern, Ryan Adams, and Peter Orbanz · 2018
Later among the works it cites.
A tight excess risk bound via a unified PAC-Bayesian-Rademacher-Shtarkov-MDL complexity
P.. Grünwald and N. Mehta · 2019
Closest in time.
Improved mdl estimators using local exponential family bundles applied to mixture families
Kohei Miyamoto, Andrew R. Barron, and J. Takeuchi · 2019
Closest in time.