Fetching the paper…
Reading the bibliography…
How do we compare between hypotheses that are entirely consistent with observations? The marginal likelihood (aka Bayesian evidence), which represents the probability of generating our observations from a prior, provides a distinctive approach to this foundational question, automatically encoding Occam's razor.
The theory of probability
Harold Jeffreys · 1939
Earlier work this paper cites.
The proof and measurement of association between two things
Charles Spearman · 1961
Earlier work this paper cites.
Corroboration, explanation, evolving probability, simplicity and a sharpened razor
Irving John Good · 1968
Earlier work this paper cites.
A new look at the statistical model identification
Hirotugu Akaike · 1974
Earlier work this paper cites.
Explicativity: a mathematical theory of explanation with statistical applications
IJ Good · 1977
Earlier work this paper cites.
Estimating the dimension of a model
Gideon Schwarz · 1978
Earlier work this paper cites.
Inference, method, and decision: Towards a bayesian philosophy of science, 1979
Edwin T Jaynes · 1979
Earlier work this paper cites.
Bayes factors and choice criteria for linear models
Adrian FM Smith and David J Spiegelhalter · 1980
Earlier work this paper cites.
Bayesian inductive inference and maximum entropy
Stephen F Gull · 1988
Earlier work this paper cites.
From laplace to supernova sn 1987a: Bayesian inference in astrophysics
Thomas J Loredo · 1990
Earlier work this paper cites.
Minimal bayesian testing of precise hypotheses, model selection, and ockham’s razor
JO Berger and WH Jeffreys · 1991
Earlier work this paper cites.
Sharpening ockham’s razor on a bayesian strop
William H Jefferys and James O Berger · 1991
Earlier work this paper cites.
Structural risk minimization for character recognition
Isabelle Guyon, Vladimir Vapnik, Bernhard Boser, Leon Bottou, and Sara A Solla · 1992
Earlier work this paper cites.
The evidence framework applied to classification networks
David J. C. MacKay · 1992
Earlier work this paper cites.
Bayes factors
Robert E Kass and Adrian E Raftery · 1995
Earlier work this paper cites.
Probable networks and plausible predictions?a review of practical Bayesian methods for supervised neural networks
David JC MacKay · 1995
Earlier work this paper cites.
The intrinsic Bayes factor for model selection and prediction
James O Berger and Luis R Pericchi · 1996
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Some pac-bayesian theorems
David A McAllester · 1998
Earlier work this paper cites.
The role of occam’s razor in knowledge discovery
Pedro Domingos · 1999
Earlier work this paper cites.
Pac-bayesian model averaging
David A McAllester · 1999
Earlier work this paper cites.
Bounds for averaging classifiers
John Langford and Matthias Seeger · 2001
Earlier work this paper cites.
Automatic choice of dimensionality for pca
Thomas P Minka · 2001
Earlier work this paper cites.
Annealed importance sampling
Radford M Neal · 2001
Earlier work this paper cites.
Occam’s razor
Carl Edward Rasmussen and Zoubin Ghahramani · 2001
Earlier work this paper cites.
Bayesian measures of model complexity and fit
David J Spiegelhalter, Nicola G Best, Bradley P Carlin, and Angelika Van Der Linde · 2002
Earlier work this paper cites.
Information theory, inference and learning algorithms
David JC MacKay · 2003
Earlier work this paper cites.
A note on the pac bayesian theorem
Andreas Maurer · 2004
Earlier work this paper cites.
On the benefits of invariance in neural networks
Clare Lyle, Mark van der Wilk, Marta Kwiatkowska, Yarin Gal, and Benjamin Bloem-Reddy · 2005
Earlier work this paper cites.
Gaussian processes for Machine Learning
C. E. Rasmussen and C. K. I. Williams · 2006
Earlier work this paper cites.
Pac-bayesian supervised classification: the thermodynamics of statistical learning
Olivier Catoni · 2007
Cited alongside, same era.
Gaussian processes for machine learning (GPML) toolbox
Carl Edward Rasmussen and Hannes Nickisch · 2010
Cited alongside, same era.
Asymptotic equivalence of bayes cross validation and widely applicable information criterion in singular learning theory
Sumio Watanabe and Manfred Opper · 2010
Cited alongside, same era.
David MacKay and Occam’s Razor
Andrew Gelman · 2011
Cited alongside, same era.
Bayesian learning via stochastic gradient Langevin dynamics
Max Welling and Yee W Teh · 2011
Cited alongside, same era.
Bayesian data analysis
Andrew Gelman, John B Carlin, Hal S Stern, David B Dunson, Aki Vehtari, and Donald B Rubin · 2013
Fixing a broken elbo
Alexander Alemi, Ben Poole, Ian Fischer, Joshua Dillon, Rif A Saurous, and Kevin Murphy · 2018
Later among the works it cites.
Foundations of machine learning
Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar · 2018
Later among the works it cites.
A scalable Laplace approximation for neural networks
Hippolyt Ritter, Aleksandar Botev, and David Barber · 2018
Later among the works it cites.
Learning invariances using the marginal likelihood
Mark van der Wilk, Matthias Bauer, ST John, and James Hensman · 2018
Later among the works it cites.
Non-vacuous generalization bounds at the imagenet scale: a pac-bayesian compression approach
Wenda Zhou, Victor Veitch, Morgane Austern, Ryan P Adams, and Peter Orbanz · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Gaussian processes for big data
J Hensman, N Fusi, and N.D. Lawrence · 2013
Cited alongside, same era.
Stochastic variational inference
Matthew D Hoffman, David M Blei, Chong Wang, and John Paisley · 2013
Cited alongside, same era.
Auto-encoding variational Bayes
Diederik P Kingma and Max Welling · 2013
Cited alongside, same era.
Gaussian process kernels for pattern discovery and extrapolation
Andrew Gordon Wilson and Ryan Prescott Adams · 2013
Cited alongside, same era.
Understanding predictive information criteria for bayesian models
Andrew Gelman, Jessica Hwang, and Aki Vehtari · 2014
Cited alongside, same era.
Auto-encoding variational Bayes
Diederik P Kingma and Max Welling · 2014
Cited alongside, same era.
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 2019
Later among the works it cites.
Fixup initialization: Residual learning without normalization
Hongyi Zhang, Yann N Dauphin, and Tengyu Ma · 2019
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Later among the works it cites.
On the marginal likelihood and cross-validation
Edwin Fong and Chris C Holmes · 2020
Later among the works it cites.
Marginal likelihood computation for model selection and hypothesis testing: an extensive review
Fernando Llorente, Luca Martino, David Delgado, and Javier Lopez-Santiago · 2020
Later among the works it cites.
Rethinking parameter counting in deep models: Effective dimensionality revisited
Wesley J Maddox, Gregory Benton, and Andrew Gordon Wilson · 2020
Later among the works it cites.
Bayesian meta-learning for the few-shot setting via deep kernels
Massimiliano Patacchiola, Jack Turner, Elliot J Crowley, Michael O’Boyle, and Amos Storkey · 2020
Later among the works it cites.
Marginal likelihood gradient for bayesian neural networks
Marcin B Tomczak and Richard E Turner · 2020
Later among the works it cites.
How good is the Bayes posterior in deep neural networks really?
Florian Wenzel, Kevin Roth, Bastiaan S Veeling, Jakub Światkowski, Linh Tran, Stephan Mandt, Jasper Snoek, Tim Salimans, Rodolphe Jenatton, and Sebastian Nowozin · 2020
Later among the works it cites.
Bayesian deep learning and a probabilistic perspective of generalization
Andrew Gordon Wilson and Pavel Izmailov · 2020
Later among the works it cites.
User-friendly introduction to PAC-Bayes bounds
Pierre Alquier · 2021
Later among the works it cites.
Laplace redux-effortless bayesian deep learning
Erik Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen, Matthias Bauer, and Philipp Hennig · 2021
Later among the works it cites.
On the role of data in pac-bayes bounds
Gintare Karolina Dziugaite, Kyle Hsu, Waseem Gharbieh, Gabriel Arpino, and Daniel Roy · 2021
Later among the works it cites.
Bayesian neural network priors revisited
Vincent Fortuin, Adrià Garriga-Alonso, Florian Wenzel, Gunnar Rätsch, Richard Turner, Mark van der Wilk, and Laurence Aitchison · 2021
Later among the works it cites.
Scalable marginal likelihood estimation for model selection in deep learning
Alexander Immer, Matthias Bauer, Vincent Fortuin, Gunnar Rätsch, and Mohammad Emtiyaz Khan · 2021
Later among the works it cites.
What are Bayesian neural network posteriors really like?
Pavel Izmailov, Sharad Vikram, Matthew D Hoffman, and Andrew Gordon Wilson · 2021
Later among the works it cites.
The promises and pitfalls of deep kernel learning
Sebastian W Ober, Carl E Rasmussen, and Mark van der Wilk · 2021
Later among the works it cites.
From predictions to decisions: The importance of joint predictive distributions
Zheng Wen, Ian Osband, Chao Qin, Xiuyuan Lu, Morteza Ibrahimi, Vikranth Dwaracherla, Mohammad Asghari, and Benjamin Van Roy · 2021
Later among the works it cites.
Differentiable annealed importance sampling and the perils of gradient noise
Guodong Zhang, Kyle Hsu, Jianing Li, Chelsea Finn, and Roger B Grosse · 2021
Later among the works it cites.
Adapting the linearised laplace model evidence for modern deep learning
Javier Antorán, David Janz, James U Allingham, Erik Daxberger, Riccardo Rb Barbano, Eric Nalisnick, and José Miguel Hernández-Lobato · 2022
Closest in time.
A pac-bayesian generalization bound for equivariant networks
Arash Behboodi, Gabriele Cesa, and Taco Cohen · 2022
Closest in time.
On uncertainty, tempering, and data augmentation in bayesian classification
Sanyam Kapoor, Wesley J Maddox, Pavel Izmailov, and Andrew Gordon Wilson · 2022
Closest in time.
Pac-bayes compression bounds so tight that they can explain generalization
Sanae Lotfi, Marc Finzi, Sanyam Kapoor, Andres Potapczynski, Micah Goldblum, and Andrew Gordon Wilson · 2022
Closest in time.
Pacm-bayes: Narrowing the empirical risk gap in the misspecified bayesian regime
Warren R Morningstar, Alex Alemi, and Joshua V Dillon · 2022
Closest in time.
Last layer marginal likelihood for invariance learning
Pola Schwöbel, Martin Jørgensen, Sebastian W Ober, and Mark Van Der Wilk · 2022
Closest in time.