Fetching the paper…
Reading the bibliography…
In recent years, the transformer has established itself as a workhorse in many applications ranging from natural language processing to reinforcement learning.
Policy gradient methods for reinforcement learning with function approximation
Richard Sutton, David Mcallester, Satinder Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Mnist handwritten digit database
Yann LeCun, Corinna Cortes, and CJ Burges · 2010
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Max Welling and Yee Whye Teh · 2011
Earlier work this paper cites.
Mcmc using hamiltonian dynamics
Radford Neal · 2012
Earlier work this paper cites.
Stochastic gradient hamiltonian monte carlo, 2014
Tianqi Chen, Emily B. Fox, and Carlos Guestrin · 2014
Earlier work this paper cites.
Weight uncertainty in neural networks, 2015
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra · 2015
Earlier work this paper cites.
Variational dropout and the local reparameterization trick, 2015
Diederik P. Kingma, Tim Salimans, and Max Welling · 2015
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning, 2016
Yarin Gal and Zoubin Ghahramani · 2016
Earlier work this paper cites.
Structured and efficient variational deep learning with matrix gaussian posteriors, 2016
Christos Louizos and Max Welling · 2016
Earlier work this paper cites.
Concrete dropout, 2017
Yarin Gal, Jiri Hron, and Alex Kendall · 2017
Earlier work this paper cites.
On calibration of modern neural networks, 2017
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger · 2017
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2017
Diederik P. Kingma and Jimmy Ba · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles, 2017
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Earlier work this paper cites.
Multiplicative normalizing flows for variational bayesian neural networks, 2017
Christos Louizos and Max Welling · 2017
Earlier work this paper cites.
Attention is all you need, 2017
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Variational attention for sequence-to-sequence models, 2018
Hareesh Bahuleyan, Lili Mou, Olga Vechtomova, and Pascal Poupart · 2018
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Earlier work this paper cites.
Latent alignment and variational attention, 2018
Yuntian Deng, Yoon Kim, Justin Chiu, Demi Guo, and Alexander M. Rush · 2018
Earlier work this paper cites.
Music transformer, 2018
Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Noam Shazeer, Ian Simon, Curtis Hawthorne, Andrew M. Dai, Matthew D. Hoffman, Monica Dinculescu, and Douglas Eck · 2018
Earlier work this paper cites.
Bayesian inference with anchored ensembles of neural networks, and application to exploration in reinforcement learning, 2018
Tim Pearce, Nicolas Anastassacos, Mohamed Zaki, and Andy Neely · 2018
Earlier work this paper cites.
A scalable laplace approximation for neural networks
Hippolyt Ritter, Aleksandar Botev, and David Barber · 2018
Cited alongside, same era.
Flipout: Efficient pseudo-independent weight perturbations on mini-batches, 2018
Yeming Wen, Paul Vicol, Jimmy Ba, Dustin Tran, and Roger Grosse · 2018
Cited alongside, same era.
Conservative uncertainty estimation by fitting prior networks
Kamil Ciosek, Vincent Fortuin, Ryota Tomioka, Katja Hofmann, and Richard Turner · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Implicit reparameterization gradients, 2019
Michael Figurnov, Shakir Mohamed, and Andriy Mnih · 2019
Cited alongside, same era.
Dirichlet variational autoencoder, 2019
Weonyoung Joo, Wonsung Lee, Sungrae Park, and Il-Chul Moon · 2019
The monte carlo transformer: a stochastic self-attention model for sequence prediction, 2020
Alice Martin, Charles Ollion, Florian Strub, Sylvain Le Corff, and Olivier Pietquin · 2020
Later among the works it cites.
Uncertainty in neural networks: Approximately bayesian ensembling, 2020
Tim Pearce, Felix Leibfried, Alexandra Brintrup, Mohamed Zaki, and Andy Neely · 2020
Later among the works it cites.
The k-tied normal distribution: A compact parameterization of gaussian mean field posteriors in bayesian neural networks, 2020
Jakub Swiatkowski, Kevin Roth, Bastiaan S. Veeling, Linh Tran, Joshua V. Dillon, Jasper Snoek, Stephan Mandt, Tim Salimans, Rodolphe Jenatton, and Sebastian Nowozin · 2020
Later among the works it cites.
Efficient low rank gaussian variational inference for neural networks
Marcin Tomczak, Siddharth Swaroop, and Richard Turner · 2020
Later among the works it cites.
How good is the bayes posterior in deep neural networks really?
Florian Wenzel, Kevin Roth, Bastiaan S Veeling, Jakub Świątkowski, Linh Tran, Stephan Mandt, Jasper Snoek, Tim Salimans, Rodolphe Jenatton, and Sebastian Nowozin · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Slang: Fast structured covariance approximations for bayesian deep learning with natural gradient, 2019
Aaron Mishkin, Frederik Kunstner, Didrik Nielsen, Mark Schmidt, and Mohammad Emtiyaz Khan · 2019
Cited alongside, same era.
On the validity of bayesian neural networks for uncertainty estimation, 2019
John Mitros and Brian Mac Namee · 2019
Cited alongside, same era.
"musenet.", 2019
OpenAI Payne, Christine · 2019
Cited alongside, same era.
Bayesian layers: A module for neural network uncertainty, 2019
Dustin Tran, Michael W. Dusenberry, Mark van der Wilk, and Danijar Hafner · 2019
Cited alongside, same era.
Cyclical stochastic gradient mcmc for bayesian deep learning
Ruqi Zhang, Chunyuan Li, Jianyi Zhang, Changyou Chen, and Andrew Gordon Wilson · 2019
Cited alongside, same era.
Repulsive attention: Rethinking multi-head attention as bayesian inference, 2020
Bang An, Jie Lyu, Zhenyi Wang, Chunyuan Li, Changwei Hu, Fei Tan, Ruiyi Zhang, Yifan Hu, and Changyou Chen · 2020
Cited alongside, same era.
Later among the works it cites.
Bayesian deep learning and a probabilistic perspective of generalization
Andrew Gordon Wilson and Pavel Izmailov · 2020
Later among the works it cites.
Visual transformers: Token-based image representation and processing for computer vision, 2020
Bichen Wu, Chenfeng Xu, Xiaoliang Dai, Alvin Wan, Peizhao Zhang, Zhicheng Yan, Masayoshi Tomizuka, Joseph Gonzalez, Kurt Keutzer, and Peter Vajda · 2020
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling, 2021
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Michael Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch · 2021
Closest in time.
Repulsive deep ensembles are bayesian
Francesco D’Angelo and Vincent Fortuin · 2021
Closest in time.
On stein variational neural network ensembles
Francesco D’Angelo, Vincent Fortuin, and Florian Wenzel · 2021
Closest in time.
Bayesian deep learning via subnetwork inference, 2021
Erik Daxberger, Eric Nalisnick, James Urquhart Allingham, Javier Antorán, and José Miguel Hernández-Lobato · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale, 2021
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Closest in time.
Priors in bayesian deep learning: A review
Vincent Fortuin · 2021
Closest in time.
Exact langevin dynamics with stochastic gradients
Adrià Garriga-Alonso and Vincent Fortuin · 2021
Closest in time.
Ast: Audio spectrogram transformer, 2021
Yuan Gong, Yu-An Chung, and James Glass · 2021
Closest in time.
What are bayesian neural network posteriors really like?
Pavel Izmailov, Sharad Vikram, Matthew D Hoffman, and Andrew Gordon Wilson · 2021
Closest in time.
Transgan: Two pure transformers can make one strong gan, and that can scale up, 2021
Yifan Jiang, Shiyu Chang, and Zhangyang Wang · 2021
Closest in time.
Segmenter: Transformer for semantic segmentation, 2021
Robin Strudel, Ricardo Garcia, Ivan Laptev, and Cordelia Schmid · 2021
Closest in time.
Bayesian transformer language models for speech recognition, 2021
Boyang Xue, Jianwei Yu, Junhao Xu, Shansong Liu, Shoukang Hu, Zi Ye, Mengzhe Geng, Xunying Liu, and Helen Meng · 2021
Closest in time.
Videogpt: Video generation using vq-vae and transformers, 2021
Wilson Yan, Yunzhi Zhang, Pieter Abbeel, and Aravind Srinivas · 2021
Closest in time.