Fetching the paper…
Reading the bibliography…
We give extensive empirical evidence against the common belief that variational learning is ineffective for large neural networks.
A practical Bayesian framework for backpropagation networks
MacKay, D. J. C · 1992
Earlier work this paper cites.
Online model selection based on the variational Bayes
Sato, M.-A · 2001
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Dolan, W. B. and Brockett, C · 2005
Earlier work this paper cites.
Automated flower classification over a large number of classes
Nilsback, M.-E. and Zisserman, A · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Practical variational inference for neural networks
Graves, A · 2011
Earlier work this paper cites.
Reading digits in natural images with unsupervised feature learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient Langevin dynamics
Welling, M. and Teh, Y. W · 2011
Earlier work this paper cites.
The winograd schema challenge
Levesque, H., Davis, E., and Morgenstern, L · 2012
Earlier work this paper cites.
Stochastic variational inference
Hoffman, M. D., Blei, D. M., Wang, C., and Paisley, J · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C · 2013
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Cho, K., Van Merriënboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Weight uncertainty in neural networks
Blundell, C., Cornebise, J., Kavukcuoglu, K., and Wierstra, D · 2015
Earlier work this paper cites.
Equilibrated adaptive learning rates for non-convex optimization
Dauphin, Y., De Vries, H., and Bengio, Y · 2015
Earlier work this paper cites.
Probabilistic backpropagation for scalable learning of Bayesian neural networks
Hernández-Lobato, J. M. and Adams, R · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Variational dropout and the local reparameterization trick
Kingma, D. P., Salimans, T., and Welling, M · 2015
Earlier work this paper cites.
Tiny ImageNet visual recognition challenge
Le, Y. and Yang, X. S · 2015
Earlier work this paper cites.
Dropout as a Bayesian approximation: Representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z · 2016
Earlier work this paper cites.
Gaussian error linear units (GELUs)
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Cer, D., Diab, M., Agirre, E., Lopez-Gazpio, I., and Specia, L · 2017
Cited alongside, same era.
Densely connected convolutional networks
Huang, G., Liu, Z., van der Maaten, L., and Weinberger, K. Q · 2017
Cited alongside, same era.
Conjugate-computation variational inference: Converting variational inference in non-conjugate models to inferences in conjugate models
Khan, M. E. and Lin, W · 2017
Cited alongside, same era.
Understanding black-box predictions via influence functions
Koh, P. W. and Liang, P · 2017
Cited alongside, same era.
Simple and scalable predictive uncertainty estimation using deep ensembles
Lakshminarayanan, B., Pritzel, A., and Blundell, C · 2017
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Later among the works it cites.
Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift
Snoek, J., Ovadia, Y., Fertig, E., Lakshminarayanan, B., Nowozin, S., Sculley, D., Dillon, J. V., Ren, J., and Nado, Z · 2019
Later among the works it cites.
On the expressiveness of approximate inference in Bayesian neural networks
Foong, A., Burt, D., Li, Y., and Turner, R · 2020
Later among the works it cites.
Handling the positive-definite constraint in the Bayesian learning rule
Lin, W., Schmidt, M., and Khan, M. E · 2020
Later among the works it cites.
How good is the bayes posterior in deep neural networks really?
Wenzel, F., Roth, K., Veeling, B. S., Swiatkowski, J., Tran, L., Mandt, S., Snoek, J., Salimans, T., Jenatton, R., and Nowozin, S · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Loshchilov, I. and Hutter, F · 2017
Cited alongside, same era.
Overpruning in variational Bayesian neural networks
Trippe, B. and Turner, R · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
Fast and scalable Bayesian deep learning by weight-perturbation in Adam
Khan, M. E., Nielsen, D., Tangkaratt, V., Lin, W., Gal, Y., and Srivastava, A · 2018
Cited alongside, same era.
Enhancing the reliability of out-of-distribution image detection in neural networks
Liang, S., Li, Y., and Srikant, R · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2018
Cited alongside, same era.
Neural network acceptability judgments
Warstadt, A., Singh, A., and Bowman, S. R · 2018
Cited alongside, same era.
Basu, S., Pope, P., and Feizi, S · 2021
Later among the works it cites.
Laplace redux – effortless Bayesian deep learning
Daxberger, E., Kristiadi, A., Immer, A., Eschenhagen, R., Bauer, M., and Hennig, P · 2021
Later among the works it cites.
SAE: Sequential anchored ensembles
Delaunoy, A. and Louppe, G · 2021
Later among the works it cites.
Sharpness-aware minimization for efficiently improving generalization
Foret, P., Kleiner, A., Mobahi, H., and Neyshabur, B · 2021
Later among the works it cites.
What are Bayesian neural network posteriors really like?
Izmailov, P., Vikram, S., Hoffman, M. D., and Wilson, A. G · 2021
Later among the works it cites.
Khan, M. E. and Rue, H · 2021
Later among the works it cites.
Disentangling the roles of curation, data-augmentation and the prior in the cold posterior effect
Noci, L., Roth, K., Bachmann, G., Nowozin, S., and Hofmann, T · 2021
Later among the works it cites.
AdaHessian: an adaptive second order optimizer for machine learning
Yao, Z., Gholami, A., Shen, S., Mustafa, M., Keutzer, K., and Mahoney, M. W · 2021
Later among the works it cites.
Wide mean-field Bayesian neural networks ignore the data
Coker, B., Bruinsma, W. P., Burt, D. R., Pan, W., and Doshi-Velez, F · 2022
Later among the works it cites.
Bayesian neural network priors revisited
Fortuin, V., Garriga-Alonso, A., Wenzel, F., Rätsch, G., Turner, R., van der Wilk, M., and Aitchison, L · 2022
Later among the works it cites.
An empirical analysis of compute-optimal large language model training
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Vinyals, O., Rae, J. W., and Sifre, L · 2022
Later among the works it cites.
Evaluating approximate inference in Bayesian deep learning
Wilson, A. G., Izmailov, P., Hoffman, M. D., Gal, Y., Li, Y., Pradier, M. F., Vikram, S., Foong, A., Lotfi, S., and Farquhar, S · 2022
Later among the works it cites.
DeBERTav3: Improving deBERTa using ELECTRA-style pre-training with gradient-disentangled embedding sharing
He, P., Gao, J., and Chen, W · 2023
Later among the works it cites.
SAM as an optimal relaxation of Bayes
Möllenhoff, T. and Khan, M. E · 2023
Later among the works it cites.
The memory perturbation equation: Understanding model’s sensitivity to data
Nickl, P., Xu, L., Tailor, D., Möllenhoff, T., and Khan, M. E · 2023
Later among the works it cites.
Model merging by uncertainty-based gradient matching
Daheim, N., Möllenhoff, T., Ponti, E. M., Gurevych, I., and Khan, M. E · 2024
Closest in time.
Sophia: A scalable stochastic second-order optimizer for language model pre-training
Liu, H., Li, Z., Hall, D., Liang, P., and Ma, T · 2024
Closest in time.