Fetching the paper…
Reading the bibliography…
We introduce Attention Free Transformer (AFT), an efficient variant of Transformers that eliminates the need for dot product self attention.
Large text compression benchmark, 2011
Matt Mahoney · 2011
Earlier work this paper cites.
Deep residual learning for image recognition, 2015
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax, 2017
Eric Jang, Shixiang Gu, and Ben Poole · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Łukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran · 2018
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Earlier work this paper cites.
Ccnet: Criss-cross attention for semantic segmentation
Zilong Huang, Xinggang Wang, Lichao Huang, C. Huang, Yunchao Wei, and Wenyu Liu · 2019
Earlier work this paper cites.
Asymmetric non-local neural networks for semantic segmentation
Zhen Zhu, Mengdu Xu, Song Bai, Tengteng Huang, and X. Bai · 2019
Earlier work this paper cites.
Interlaced sparse self-attention for semantic segmentation
Lang Huang, Y. Yuan, Jianyuan Guo, Chao Zhang, X. Chen, and Jingdong Wang · 2019
Earlier work this paper cites.
Stand-alone self-attention in vision models
Prajit Ramachandran, Niki Parmar, Ashish Vaswani, I. Bello, Anselm Levskaya, and Jonathon Shlens · 2019
Cited alongside, same era.
Adaptive attention span in transformers
Sainbayar Sukhbaatar, E. Grave, P. Bojanowski, and Armand Joulin · 2019
Cited alongside, same era.
Pay less attention with lightweight and dynamic convolutions
Felix Wu, Angela Fan, Alexei Baevski, Yann Dauphin, and M. Auli · 2019
Cited alongside, same era.
Decoupled weight decay regularization, 2019
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Z. Yang, Yiming Yang, J. Carbonell, Quoc V. Le, and R. Salakhutdinov · 2019
Cited alongside, same era.
Synthesizer: Rethinking self-attention in transformer models, 2020
Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, and Che Zheng · 2020
Later among the works it cites.
Rethinking attention with performers, 2020
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, David Belanger, Lucy Colwell, and Adrian Weller · 2020
Later among the works it cites.
Efficient transformers: A survey, 2020
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler · 2020
Later among the works it cites.
Axial-deeplab: Stand-alone axial-attention for panoptic segmentation
Huiyu Wang, Y. Zhu, B. Green, H. Adam, A. Yuille, and Liang-Chieh Chen · 2020
Later among the works it cites.
Efficient content-based sparse attention with routing transformers
Aurko Roy, M. Saffar, Ashish Vaswani, and David Grangier · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2020
Cited alongside, same era.
Reformer: The efficient transformer
Nikita Kitaev, L. Kaiser, and Anselm Levskaya · 2020
Cited alongside, same era.
Compressive transformers for long-range sequence modelling
Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, and T. Lillicrap · 2020
Cited alongside, same era.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma · 2020
Cited alongside, same era.
Transformers are rnns: Fast autoregressive transformers with linear attention
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret · 2020
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever
Cited in the paper.
Yi Tay, Dara Bahri, L. Yang, Donald Metzler, and D. Juan · 2020
Later among the works it cites.
Random feature attention
Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah Smith, and Lingpeng Kong · 2021
Closest in time.
Lambdanetworks: Modeling long-range interactions without attention
Irwan Bello · 2021
Closest in time.
Mlp-mixer: An all-mlp architecture for vision, 2021
Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, and Alexey Dosovitskiy · 2021
Closest in time.
Pay attention to mlps, 2021
Hanxiao Liu, Zihang Dai, David R. So, and Quoc V. Le · 2021
Closest in time.