Fetching the paper…
Reading the bibliography…
While Transformer networks benefit from a global receptive field, their quadratic cost relative to sequence length restricts their application to long sequences and high-resolution inputs.
A hierarchical O(NlogN) force-calculation algorithm
J. Barnes and P. Hut · 1986
Earlier work this paper cites.
A fast algorithm for particle simulations
L. Greengard and V. Rokhlin · 1987
Earlier work this paper cites.
A short course on fast multipole methods
R. Beatson and L. Greengard · 1997
Earlier work this paper cites.
A sparse matrix arithmetic based on H-matrices. part I: Introduction to H-matrices
W. Hackbusch · 1999
Earlier work this paper cites.
Matrices with Hierarchical Low-Rank Structures
M. Bebendorf · 2008
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2015
Earlier work this paper cites.
Fast multipole methods
P.-G. Martinsson · 2015
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
TVM: An automated end-to-end optimizing compiler for deep learning
T. Chen, T. Moreau, Z. Jiang, L. Zheng, E. Yan, M. Cowan, H. Shen, L. Wang, Y. Hu, L. Ceze, C. Guestrin, and A. Krishnamurthy · 2018
Earlier work this paper cites.
Generating Wikipedia by summarizing long sequences
P. J. Liu, M. Saleh, E. Pot, B. Goodrich, R. Sepassi, L. Kaiser, and N. Shazeer · 2018
Earlier work this paper cites.
Generating long sequences with sparse transformers
R. Child, S. Gray, A. Radford, and I. Sutskever · 2019
Earlier work this paper cites.
Axial attention in multidimensional transformers
J. Ho and N. Kalchbrenner · 2019
Earlier work this paper cites.
Music transformer
C.-Z. A. Huang, A. Vaswani, J. Uszkoreit, I. Simon, C. Hawthorne, N. Shazeer, A. M. Dai, M. D. Hoffman, M. Dinculescu, and D. Eck · 2019
Earlier work this paper cites.
fairseq: A fast, extensible toolkit for sequence modeling
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli · 2019
Earlier work this paper cites.
ETC: Encoding long and structured inputs in transformers
J. Ainslie, S. Ontanon, C. Alberti, V. Cvicek, Z. Fisher, P. Pham, A. Ravula, S. Sanghai, Q. Wang, and L. Yang · 2020
Cited alongside, same era.
Longformer: The long-document transformer
I. Beltagy, M. E. Peters, and A. Cohan · 2020
Cited alongside, same era.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Cited alongside, same era.
Conformer: Convolution-augmented Transformer for Speech Recognition
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang · 2020
Cited alongside, same era.
Transformers are RNNs: Fast autoregressive transformers with linear attention
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret · 2020
Swin transformer: Hierarchical vision transformer using shifted windows
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo · 2021
Later among the works it cites.
Random feature attention
H. Peng, N. Pappas, D. Yogatama, R. Schwartz, N. Smith, and L. Kong · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Later among the works it cites.
Efficient content-based sparse attention with routing transformers
A. Roy, M. Saffar, A. Vaswani, and D. Grangier · 2021
Later among the works it cites.
Cluster-Former: Clustering-based sparse transformer for question answering
S. Wang, L. Zhou, Z. Gan, Y.-C. Chen, Y. Fang, S. Sun, Y. Cheng, and J. Liu · 2021
Later among the works it cites.
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Reformer: The efficient transformer
N. Kitaev, L. Kaiser, and A. Levskaya · 2020
Cited alongside, same era.
Fast transformers with clustered attention
A. Vyas, A. Katharopoulos, and F. Fleuret · 2020
Cited alongside, same era.
Linformer: Self-attention with linear complexity
S. Wang, B. Z. Li, M. Khabsa, H. Fang, and H. Ma · 2020
Cited alongside, same era.
Big Bird: Transformers for longer sequences
M. Zaheer, G. Guruganesh, K. A. Dubey, J. Ainslie, C. Alberti, S. Ontanon, P. Pham, A. Ravula, Q. Wang, L. Yang, and A. Ahmed · 2020
Cited alongside, same era.
Rethinking attention with Performers
K. M. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, P. Hawkins, J. Q. Davis, A. Mohiuddin, L. Kaiser, D. B. Belanger, L. J. Colwell, and A. Weller · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Cited alongside, same era.
Halonet: Efficient attention for vision transformers
A. A. Hassani, C. Walton, A. Aitchison, M. Nunez, and S. Belongie · 2021
Cited alongside, same era.
W. Wang, E. Xie, X. Li, D. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Wang · 2021
Later among the works it cites.
Segformer: Simple and efficient design for semantic segmentation with transformers
E. Xie, W. Wang, Z. Yu, A. Anandkumar, J. M. Alvarez, and P. Luo · 2021
Later among the works it cites.
H-Transformer-1D: Fast One-Dimensional Hierarchical Attention for Sequences
Z. Zhu and R. Soricut · 2021
Later among the works it cites.
cosFormer: Rethinking softmax in attention
Z. Qin, W. Sun, H. Deng, D. Li, Y. Wei, B. Lv, J. Yan, L. Kong, and Y. Zhong · 2022
Later among the works it cites.
A generalist agent
S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-Maron, M. Giménez, Y. Sulsky, J. Kay, J. T. Springenberg, T. Eccles, J. Bruce, A. Razavi, A. Edwards, N. Heess, Y. Chen, R. Hadsell, O. Vinyals, M. Bordbar, and N. de Freitas · 2022
Later among the works it cites.
Efficient transformers: A survey
Y. Tay, M. Dehghani, D. Bahri, and D. Metzler · 2022
Later among the works it cites.
ClusterFormer: Neural clustering attention for efficient and effective transformer
N. Wang, G. Gan, P. Zhang, S. Zhang, J. Wei, Q. Liu, and X. Jiang · 2022
Later among the works it cites.
Multi resolution analysis (MRA) for approximate self-attention
Z. Zeng, S. Pal, J. Kline, G. M. Fung, and V. Singh · 2022
Later among the works it cites.