Fetching the paper…
Reading the bibliography…
Attentional mechanisms are order-invariant.
Backpropagation applied to handwritten zip code recognition
Y. LeCun, B. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. Hubbard, and L. D. Jackel · 1989
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning
Ali Rahimi and Benjamin Recht · 2008
Earlier work this paper cites.
Compact random feature maps
Raffay Hamid, Ying Xiao, Alex Gittens, and Dennis DeCoste · 2014
Earlier work this paper cites.
Fastfood: Approximate kernel expansions in loglinear time
Quoc Le, Tamas Sarlos, and Alexander Smola · 2014
Earlier work this paper cites.
Microsoft COCO: common objects in context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
Quasi-monte carlo feature maps for shift-invariant kernels
Jiyan Yang, Vikas Sindhwani, Haim Avron, and Michael Mahoney · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate, 2014
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Thang Luong, Hieu Pham, and Christopher D. Manning · 2015
Earlier work this paper cites.
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin · 2017
Earlier work this paper cites.
Taming the waves: sine as activation function in deep neural networks
Giambattista Parascandolo, Heikki Huttunen, and Tuomas Virtanen · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
An improved relative self-attention mechanism for transformer with application to music generation
Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Noam Shazeer, Curtis Hawthorne, Andrew M. Dai, Matthew D. Hoffman, and Douglas Eck · 2018
Earlier work this paper cites.
Image transformer
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran · 2018
Earlier work this paper cites.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2018
Cited alongside, same era.
Fundamentals of recurrent neural network (RNN) and long short-term memory (LSTM) network
Alex Sherstinsky · 2018
Cited alongside, same era.
But how does it work in theory? linear svm with random features
Yitong Sun, Anna Gilbert, and Ambuj Tewari · 2018
Cited alongside, same era.
Attention augmented convolutional networks
Irwan Bello, Barret Zoph, Ashish Vaswani, Jonathon Shlens, and Quoc V. Le · 2019
Cited alongside, same era.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Cited alongside, same era.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Later among the works it cites.
Masked language modeling for proteins via linearly scalable long-context transformers, 2020
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, David Belanger, Lucy Colwell, and Adrian Weller · 2020
Later among the works it cites.
Reformer: The efficient transformer
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya · 2020
Later among the works it cites.
Mapping natural language instructions to mobile UI action sequences
Yang Li, Jiacong He, Xin Zhou, Yuan Zhang, and Jason Baldridge · 2020
Later among the works it cites.
Widget captioning: Generating natural language description for mobile user interface elements
Yang Li, Gang Li, Luheng He, Jingjie Zheng, Hong Li, and Zhiwei Guan · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Learning adaptive random features
Yanjun Li, Kai Zhang, Jun Wang, and Sanjiv Kumar · 2019
Cited alongside, same era.
On the relation between position information and sentence length in neural machine translation
Masato Neishi and Naoki Yoshinaga · 2019
Cited alongside, same era.
Recurrent space-time graph neural networks
Andrei Nicolicioiu, Iulia Duta, and Marius Leordeanu · 2019
Cited alongside, same era.
Stand-alone self-attention in vision models
Prajit Ramachandran, Niki Parmar, Ashish Vaswani, Irwan Bello, Anselm Levskaya, and Jon Shlens · 2019
Cited alongside, same era.
Novel positional encodings to enable tree-based transformers
Vighnesh Leonardo Shiv and Chris Quirk · 2019
Cited alongside, same era.
Learning to encode position for transformer with continuous dynamical model
Xuanqing Liu, Hsiang-Fu Yu, I. Dhillon, and Cho-Jui Hsieh · 2020
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2020
Later among the works it cites.
Encoding word order in complex embeddings
Benyou Wang, Donghao Zhao, Christina Lioma, Qiuchi Li, Peng Zhang, and Jakob Grue Simonsen · 2020
Later among the works it cites.
Do we really need explicit position encodings for vision transformers?
Xiangxiang Chu, Bo Zhang, Zhi Tian, Xiaolin Wei, and Huaxia Xia · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Closest in time.
Levit: a vision transformer in convnet’s clothing for faster inference
Benjamin Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock, Armand Joulin, Hervé Jégou, and Matthijs Douze · 2021
Closest in time.
Deberta: Decoding-enhanced bert with disentangled attention, 2021
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen · 2021
Closest in time.
Relative positional encoding for transformers with linear complexity
Antoine Liutkus, Ondřej Cífka, Shih-Lun Wu, Umut Simsekli, Yi-Hsuan Yang, and Gael Richard · 2021
Closest in time.
On position embeddings in {bert}
Benyou Wang, Lifeng Shang, Christina Lioma, Xin Jiang, Hao Yang, Qun Liu, and Jakob Grue Simonsen · 2021
Closest in time.