Fetching the paper…
Reading the bibliography…
Recently, Transformer networks have redefined the state of the art in many NLP tasks.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, and Phillipp Koehn · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Minh-Thang Luong, Hieu Pham, and Christopher D. Manning · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Xnli: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel R. Bowman, Holger Schwenk, and Veselin Stoyanov · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional Transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
On controllable sparse alternatives to softmax
Anirban Laha, Saneem Ahmed Chemmengath, Priyanka Agrawal, Mitesh Khapra, Karthik Sankaranarayanan, and Harish G Ramaswamy · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Earlier work this paper cites.
Tensor2tensor for neural machine translation
Ashish Vaswani, Samy Bengio, Eugene Brevdo, Francois Chollet, Aidan N Gomez, Stephan Gouws, Llion Jones, Łukasz Kaiser, Nal Kalchbrenner, Niki Parmar, et al · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Earlier work this paper cites.
On identifiability in Transformers
Gino Brunner, Yang Liu, Damián Pascual, Oliver Richter, Massimiliano Ciaramita, and Roger Wattenhofer · 2019
Earlier work this paper cites.
Generating long sequences with sparse Transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Cited alongside, same era.
Adaptively sparse Transformers
Gonçalo M Correia, Vlad Niculae, and André FT Martins · 2019
Cited alongside, same era.
Fine-tune BERT with sparse self-attention mechanism
Baiyun Cui, Yingming Li, Ming Chen, and Zhongfei Zhang · 2019
Cited alongside, same era.
Qipeng Guo, Xipeng Qiu, Pengfei Liu, Yunfan Shao, Xiangyang Xue, and Zheng Zhang · 2019
Cited alongside, same era.
Deep, skinny neural networks are not universal approximators
Jesse Johnson · 2019
Cited alongside, same era.
Bp-Transformer: Modelling long-range context via binary partitioning
Zihao Ye, Qipeng Guo, Quan Gan, Xipeng Qiu, and Zheng Zhang · 2019
Later among the works it cites.
Explicit sparse Transformer: Concentrated attention through explicit selection
Guangxiang Zhao, Junyang Lin, Zhiyuan Zhang, Xuancheng Ren, Qi Su, and Xu Sun · 2019
Later among the works it cites.
Longformer: The long-document Transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan · 2020
Closest in time.
Low-rank bottleneck in multi-head attention models
Srinadh Bhojanapalli, Chulhee Yun, Ankit Singh Rawat, Sashank J Reddi, and Sanjiv Kumar · 2020
Closest in time.
Theoretical limitations of self-attention in neural sequence models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
On the Turing completeness of modern neural network architectures
Jorge Pérez, Javier Marinković, and Pablo Barceló · 2019
Cited alongside, same era.
Sparse sequence-to-sequence models
Ben Peters, Vlad Niculae, and André FT Martins · 2019
Cited alongside, same era.
Blockwise self-attention for long document understanding
Jiezhong Qiu, Hao Ma, Omer Levy, Scott Wen-tau Yih, Sinong Wang, and Jie Tang · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Cited alongside, same era.
Adaptive attention span in Transformers
Sainbayar Sukhbaatar, Edouard Grave, Piotr Bojanowski, and Armand Joulin · 2019
Cited alongside, same era.
XLNet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le · 2019
Cited alongside, same era.
Michael Hahn · 2020
Closest in time.
Reformer: The efficient Transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya · 2020
Closest in time.
Sac: Accelerating and structuring self-attention via sparse adaptive connection
Xiaoya Li, Yuxian Meng, Qinghong Han, Fei Wu, and Jiwei Li · 2020
Closest in time.
Minimum width for universal approximation
Sejun Park, Chulhee Yun, Jaeho Lee, and Jinwoo Shin · 2020
Closest in time.
Efficient content-based sparse attention with routing Transformers
Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier · 2020
Closest in time.
Efficient transformers: A survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler · 2020
Closest in time.
Are Transformers universal approximators of sequence-to-sequence functions?
Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J Reddi, and Sanjiv Kumar · 2020
Closest in time.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al · 2020
Closest in time.