Bayesian attention modules
Xinjie Fan Fan, Shujian Zhang, Bo Chen, and Mingyuan Zhou · 2020
Later among the works it cites.
On the expressiveness of approximate inference in bayesian neural networks
Andrew Y. K. Foong, David Burt, Yingzhen Li, and Richard Turner · 2020
Later among the works it cites.
Being bayesian, even just a bit, fixes overconfidence in relu networks
Agustinus Kristiadi, Matthias Hein, and Philipp Hennig · 2020
Later among the works it cites.
Simple and principled uncertainty estimation with deterministic deep learning via distance awareness
Jeremiah Zhe Liu, Zi Lin, Shreyas Padhy, Dustin Tran, Tania Bedrax-Weiss, and Balaji Lakshminarayanan · 2020
Later among the works it cites.
Calibrating deep neural networks using focal loss
Jishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz, Philip HS Torr, and Puneet K Dokania · 2020
Later among the works it cites.
Efficiently sampling functions from gaussian process posteriors
James T. Wilson, Viacheslav Borovitskiy, Alexander Terenin, Peter Mostowsky, and Marc Peter Deisenroth · 2020
Later among the works it cites.
Pathologies in priors and inference for bayesian transformers
Tristan Cinquin, Alexander Immer, Max Horn, and Vincent Fortuin · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, and Sylvain Gelly · 2021
Later among the works it cites.
Sparse uncertainty representation in deep learning with inducing weights
Hippolyt Ritter, Martin Kukla, Cheng Zhang, and Yingzhen Li · 2021
Later among the works it cites.
Scalable variational gaussian processes via harmonic kernel decomposition
Shengyang Sun, Jiaxin Shi, A. G. Wilson, and Roger Grosse · 2021
Later among the works it cites.
Sparse within sparse gaussian processes using neighbor information
Gia-Lac Tran, Dimitrios Milios, Pietro Michiardi, and Maurizio Filippone · 2021
Later among the works it cites.
Bayesian transformer language models for speech recognition
Original
Boyang Xue, Jianwei Yu, Junhao Xu, Shansong Liu, Shoukang Hu, Zi Ye, Mengzhe Geng, Xunying Liu, and Helen Meng · 2021
Later among the works it cites.
Wide mean-field variational bayesian neural networks ignore the data
Beau Coker, Wessel P. Bruinsma, David R. Burt, Weiwei Pan, and Finale Doshi-Velez · 2022
Later among the works it cites.
Input dependent sparse gaussian processes
Bahram Jafrasteh, Carlos Villacampa-Calvo, and Daniel Hernandez-Lobato · 2022
Later among the works it cites.
Pure transformers are powerful graph learners
Original
Jinwoo Kim, Tien Dat Nguyen, Seonwoo Min, Sungjun Cho, Moontae Lee, Honglak Lee, and Seunghoon Hong · 2022
Later among the works it cites.
Deit iii: Revenge of the vit
Original
Hugo Touvron, Matthieu Cord, and Herve Jegou · 2022
Later among the works it cites.