Fetching the paper…
Reading the bibliography…
Marginal-likelihood based model-selection, even though promising, is rarely used in deep learning due to estimation difficulties.
Ix. on the problem of the most efficient tests of statistical hypotheses
Neyman, J. and Pearson, E. S · 1933
Earlier work this paper cites.
Rank correlation methods
Kendall, M. G · 1948
Earlier work this paper cites.
An empirical Bayes approach to statistics
Robbins, H · 1955
Earlier work this paper cites.
Occam’s razor
Blumer, A., Ehrenfeucht, A., Haussler, D., and Warmuth, M. K · 1987
Earlier work this paper cites.
Bayesian back-propagation
Buntine, W. L. and Weigend, A. S · 1991
Earlier work this paper cites.
Ockham’s razor and bayesian analysis
Jefferys, W. H. and Berger, J. O · 1992
Earlier work this paper cites.
A practical bayesian framework for backpropagation networks
MacKay, D. J · 1992
Earlier work this paper cites.
Probable networks and plausible predictions—a review of practical bayesian methods for supervised neural networks
MacKay, D. J · 1995
Earlier work this paper cites.
Bayesian Learning for Neural Networks
Neal, R. M · 1995
Earlier work this paper cites.
Gauss-newton approximation to bayesian learning
Foresee, F. D. and Hagan, M. T · 1997
Earlier work this paper cites.
Flat minima
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Matrix algebra from a statistician’s perspective, 1998
Harville, D. A · 1998
Earlier work this paper cites.
Occam’s razor
Rasmussen, C. E. and Ghahramani, Z · 2001
Earlier work this paper cites.
Soft margins for adaboost
Rätsch, G., Onoda, T., and Müller, K.-R · 2001
Earlier work this paper cites.
Information theory, inference and learning algorithms
MacKay, D. J · 2003
Earlier work this paper cites.
Pattern recognition and machine learning
Bishop, C. M · 2006
Earlier work this paper cites.
Gaussian processes for machine learning
Rasmussen, C. E. and Williams, C. K · 2006
Earlier work this paper cites.
Flexible and efficient Gaussian process models for machine learning
Snelson, E. L · 2007
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Bottou, L · 2010
Cited alongside, same era.
Deep Gaussian processes
Damianou, A. and Lawrence, N. D · 2013
Cited alongside, same era.
The divergence of the bfgs and gauss newton methods
Mascarenhas, W. F · 2014
Cited alongside, same era.
Weight uncertainty in neural networks
Blundell, C., Cornebise, J., Kavukcuoglu, K., and Wierstra, D · 2015
Cited alongside, same era.
Probabilistic backpropagation for scalable learning of bayesian neural networks
Hernández-Lobato, J. M. and Adams, R · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Cited alongside, same era.
Deepobs: A deep learning optimizer benchmark suite
Schneider, F., Balles, L., and Hennig, P · 2018
Later among the works it cites.
Learning invariances using the marginal likelihood
van der Wilk, M., Bauer, M., John, S., and Hensman, J · 2018
Later among the works it cites.
Noisy natural gradient as variational inference
Zhang, G., Sun, S., Duvenaud, D., and Grosse, R · 2018
Later among the works it cites.
Backpack: Packing more into backprop
Dangel, F., Kunstner, F., and Hennig, P · 2019
Later among the works it cites.
’in-between’uncertainty in bayesian neural networks
Foong, A. Y., Li, Y., Hernández-Lobato, J. M., and Turner, R. E · 2019
Later among the works it cites.
Fantastic generalization measures and where to find them
Jiang, Y., Neyshabur, B., Mobahi, H., Krishnan, D., and Bengio, S · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Martens, J. and Grosse, R · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Cited alongside, same era.
Wide residual networks
Zagoruyko, S. and Komodakis, N · 2016
Cited alongside, same era.
Practical Gauss-Newton optimisation for deep learning
Botev, A., Ritter, H., and Barber, D · 2017
Cited alongside, same era.
UCI machine learning repository, 2017
Dua, D. and Graff, C · 2017
Cited alongside, same era.
Later among the works it cites.
Approximate inference turns deep networks into gaussian processes
Khan, M. E. E., Immer, A., Abedi, E., and Korzepa, M · 2019
Later among the works it cites.
Limitations of the empirical fisher approximation for natural gradient descent
Kunstner, F., Hennig, P., and Balles, L · 2019
Later among the works it cites.
A simple baseline for bayesian uncertainty in deep learning
Maddox, W. J., Izmailov, P., Garipov, T., Vetrov, D. P., and Wilson, A. G · 2019
Later among the works it cites.
Practical deep learning with bayesian principles
Osawa, K., Swaroop, S., Khan, M. E. E., Jain, A., Eschenhagen, R., Turner, R. E., and Yokota, R · 2019
Later among the works it cites.
Bayesian image classification with deep convolutional gaussian processes
Dutordoir, V., van der Wilk, M., Artemev, A., and Hensman, J · 2020
Later among the works it cites.
On the marginal likelihood and cross-validation
Fong, E. and Holmes, C · 2020
Later among the works it cites.
Being bayesian, even just a bit, fixes overconfidence in relu networks
Kristiadi, A., Hein, M., and Hennig, P · 2020
Later among the works it cites.
Marginal likelihood computation for model selection and hypothesis testing: an extensive review
Llorente, F., Martino, L., Delgado, D., and Lopez-Santiago, J · 2020
Later among the works it cites.
A bayesian perspective on training speed and model selection
Lyle, C., Schut, L., Ru, R., Gal, Y., and van der Wilk, M · 2020
Later among the works it cites.
How good is the bayes posterior in deep neural networks really?
Wenzel, F., Roth, K., Veeling, B. S., Światkowski, J., Tran, L., Mandt, S., Snoek, J., Salimans, T., Jenatton, R., and Nowozin, S · 2020
Later among the works it cites.
Improving predictions of bayesian neural nets via local linearization
Immer, A., Korzepa, M., and Bauer, M · 2021
Closest in time.