Fetching the paper…
Reading the bibliography…
We propose a VAE for Transformers by developing a variational information bottleneck regulariser for Transformer embeddings.
Latent space secrets of denoising text-autoencoders
Tianxiao Shen, Jonas Mueller, Regina Barzilay, and Tommi S. Jaakkola. 2019 · 1905
Earlier work this paper cites.
Exchangeability and related topics
David J. Aldous. 1985 · 1983
Earlier work this paper cites.
Neural Networks for Pattern Recognition
Christopher M. Bishop. 1995 · 1995
Earlier work this paper cites.
Variational transformers for diverse response generation
Zhaojiang Lin, Genta Indra Winata, Peng Xu, Zihan Liu, and Pascale Fung. 2020 · 2003
Earlier work this paper cites.
Bayesian nonparametric learning: Expressive priors for intelligent systems
M.I. Jordan. 2010 · 2010
Earlier work this paper cites.
Dirichlet processes
Yee Whye Teh. 2010 · 2010
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P. Kingma and Max Welling. 2014 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Cited alongside, same era.
Stochastic gradient variational bayes for gamma approximating distributions
David A. Knowles. 2015 · 2015
Cited alongside, same era.
Improving variational autoencoders with inverse autoregressive flow
Diederik P. Kingma, Tim Salimans, Rafal Józefowicz, Xi Chen, Ilya Sutskever, and Max Welling. 2016 · 2016
Cited alongside, same era.
Deep variational information bottleneck
Alexander A. Alemi, Ian Fischer, Joshua V. Dillon, and Kevin Murphy. 2017 · 2017
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Adversarially regularized autoencoders for generating discrete structures
Junbo Jake Zhao, Yoon Kim, Kelly Zhang, Alexander M. Rush, and Yann LeCun. 2017 · 2017
Later among the works it cites.
Eval all, trust a few, do wrong to none: Comparing sentence generation models
Ondrej Cífka, Aliaksei Severyn, Enrique Alfonseca, and Katja Filippova. 2018 · 2018
Later among the works it cites.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
The unstoppable rise of computational linguistics in deep learning
James Henderson. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Deep sets
Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabás Póczos, Ruslan Salakhutdinov, and Alexander J. Smola. 2017 · 2017
Cited alongside, same era.
Le Fang, Tao Zeng, Chaochun Liu, Liefeng Bo, Wen Dong, and Changyou Chen. 2021 · 2021
Later among the works it cites.
Finetuning pretrained transformers into variational autoencoders
Seongmin Park and Jihwa Lee. 2021 · 2021
Later among the works it cites.