Fetching the paper…
Reading the bibliography…
Here, we show that in the data-rich setting where you only train on each datapoint once (or equivalently, you only train for one epoch), standard "maximum likelihood" training optimizes the true data generating process (DGP) loss, which is equivalent to the test loss.
Bayesian averaging of classifiers and the overfitting problem
Domingos, P · 2000
Earlier work this paper cites.
Preventing over-fitting during model selection via bayesian regularisation of the hyper-parameters
Cawley, G. C. and Talbot, N. L · 2007
Earlier work this paper cites.
Algebraic geometry and statistical learning theory , volume 25
Watanabe, S · 2009
Earlier work this paper cites.
Practical variational inference for neural networks
Graves, A · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Weight uncertainty in neural networks
Blundell, C., Cornebise, J., Kavukcuoglu, K., and Wierstra, D · 2015
Earlier work this paper cites.
Dropout as a Bayesian approximation: Representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Earlier work this paper cites.
Dziugaite, G. K. and Roy, D. M · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Lakshminarayanan, B., Pritzel, A., and Blundell, C · 2017
Cited alongside, same era.
Deep ensembles: A loss landscape perspective
Fort, S., Hu, H., and Lakshminarayanan, B · 2019
Cited alongside, same era.
Fantastic generalization measures and where to find them
Jiang, Y., Neyshabur, B., Mobahi, H., Krishnan, D., and Bengio, S · 2019
Cited alongside, same era.
Functional variational Bayesian neural networks
Sun, S., Zhang, G., Shi, J., and Grosse, R · 2019
Cited alongside, same era.
Bayesian filtering unifies adaptive and non-adaptive neural network optimization methods
Aitchison, L · 2020
Cited alongside, same era.
Bayesian neural network priors revisited
Fortuin, V., Garriga-Alonso, A., Ober, S. W., Wenzel, F., Ratsch, G., Turner, R. E., van der Wilk, M., and Aitchison, L · 2021
Later among the works it cites.
What are bayesian neural network posteriors really like?
Izmailov, P., Vikram, S., Hoffman, M. D., and Wilson, A. G. G · 2021
Later among the works it cites.
Global inducing point variational posteriors for Bayesian neural networks and deep Gaussian processes
Ober, S. W. and Aitchison, L · 2021
Later among the works it cites.
Regularizing neural networks via adversarial model perturbation
Zheng, Y., Zhang, R., and Mao, Y · 2021
Later among the works it cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Foret, P., Kleiner, A., Mobahi, H., and Neyshabur, B · 2020
Cited alongside, same era.
The pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al · 2020
Cited alongside, same era.
Variational laplace for bayesian neural networks
Unlu, A. and Aitchison, L · 2020
Cited alongside, same era.
Repulsive deep ensembles are bayesian
D’Angelo, F. and Fortuin, V · 2021
Cited alongside, same era.
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
Dodge, J., Sap, M., Marasović, A., Agnew, W., Ilharco, G., Groeneveld, D., Mitchell, M., and Gardner, M · 2021
Cited alongside, same era.
Folgoc, L. L., Baltatzis, V., Desai, S., Devaraj, A., Ellis, S., Manzanera, O. E. M., Nair, A., Qiu, H., Schnabel, J., and Glocker, B · 2021
Cited alongside, same era.
Bayesian reward models for LLM alignment
Yang, A. X., Robeyns, M., Coste, T., Wang, J., Bou-Ammar, H., and Aitchison, L
Cited in the paper.
Elazar, Y., Bhagia, A., Magnusson, I., Ravichander, A., Schwenk, D., Suhr, A., Walsh, P., Groeneveld, D., Soldaini, L., Singh, S., et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Lora ensembles for large language model fine-tuning
Wang, X., Aitchison, L., and Rudolph, M · 2023
Later among the works it cites.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Closest in time.
Position paper: Bayesian deep learning in the age of large-scale ai
Papamarkou, T., Skoularidou, M., Palla, K., Aitchison, L., Arbel, J., Dunson, D., Filippone, M., Fortuin, V., Hennig, P., Hubin, A., et al · 2024
Closest in time.
Variational learning is effective for large deep networks
Shen, Y., Daheim, N., Cong, B., Nickl, P., Marconi, G. M., Bazan, C., Yokota, R., Gurevych, I., Cremers, D., and Khan, M. E · 2024
Closest in time.