Fetching the paper…
Reading the bibliography…
Many efficient $\textit{approximate}$ self-attention techniques have become prevalent since the inception of the transformer architecture.
Generating long sequences with sparse transformers
R. Child, S. Gray, A. Radford, and I. Sutskever · 1904
Earlier work this paper cites.
Adaptively sparse transformers
G. M. Correia, V. Niculae, and A. F. T. Martins · 1909
Earlier work this paper cites.
Blockwise self-attention for long document understanding
J. Qiu, H. Ma, O. Levy, S. W. Yih, S. Wang, and J. Tang · 1911
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Köpf, E. Z. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala · 1912
Earlier work this paper cites.
Reformer: The efficient transformer
N. Kitaev, L. Kaiser, and A. Levskaya · 2001
Earlier work this paper cites.
Y. Tay, D. Bahri, L. Yang, D. Metzler, and D. Juan · 2002
Earlier work this paper cites.
Efficient content-based sparse attention with routing transformers
A. Roy, M. Saffar, A. Vaswani, and D. Grangier · 2003
Earlier work this paper cites.
Longformer: The long-document transformer
I. Beltagy, M. E. Peters, and A. Cohan · 2004
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2005
Earlier work this paper cites.
Synthesizer: Rethinking self-attention in transformer models
Y. Tay, D. Bahri, D. Metzler, D. Juan, Z. Zhao, and C. Zheng · 2005
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret · 2006
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
S. Wang, B. Z. Li, M. Khabsa, H. Fang, and H. Ma · 2006
Earlier work this paper cites.
Random features for large-scale kernel machines
A. Rahimi and B. Recht · 2007
Cited alongside, same era.
Big bird: Transformers for longer sequences
M. Zaheer, G. Guruganesh, A. Dubey, J. Ainslie, C. Alberti, S. Ontañón, P. Pham, A. Ravula, Q. Wang, L. Yang, and A. Ahmed · 2007
Cited alongside, same era.
Efficient transformers: A survey
Y. Tay, M. Dehghani, D. Bahri, and D. Metzler · 2009
Cited alongside, same era.
Y. Zhu, R. Kiros, R. Zemel, R. Salakhutdinov, R. Urtasun, A. Torralba, and S. Fidler · 2015
Cited alongside, same era.
Pointer sentinel mixture models
S. Merity, C. Xiong, J. Bradbury, and R. Socher · 2016
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
W. Fedus, B. Zoph, and N. Shazeer · 2021
Later among the works it cites.
Random feature attention
H. Peng, N. Pappas, D. Yogatama, R. Schwartz, N. Smith, and L. Kong · 2021
Later among the works it cites.
Nyströmformer: A nyström-based algorithm for approximating self-attention
Y. Xiong, Z. Zeng, R. Chakraborty, M. Tan, G. Fung, Y. Li, and V. Singh · 2021
Later among the works it cites.
Long-short transformer: Efficient transformers for language and vision
C. Zhu, W. Ping, C. Xiao, M. Shoeybi, T. Goldstein, A. Anandkumar, and B. Catanzaro · 2021
Later among the works it cites.
Palm: Scaling language modeling with pathways, 2022
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, P. Schuh, K. Shi, S. Tsvyashchenko, J. Maynez, A. Rao, P. Barnes, Y. Tay, N. Shazeer, V. Prabhakaran, E. Reif, N. Du, B. Hutchinson, R. Pope, J. Bradbury, J. Austin, M. Isard, G. Gur-Ari, P. Yin, T. Duke, A. Levskaya, S. Ghemawat, S. Dev, H. Michalewski, X. Garcia, V. Misra, K. Robinson, L. Fedus, D. Zhou, D. Ippolito, D. Luan, H. Lim, B. Zoph, A. Spiridonov, R. Sepassi, D. Dohan, S. Agrawal, M. Omernick, A. M. Dai, T. S. Pillai, M. Pellat, A. Lewkowycz, E. Moreira, R. Child, O. Polozov, K. Lee, Z. Zhou, X. Wang, B. Saeta, M. Diaz, O. Firat, M. Catasta, J. Wei, K. Meier-Hellstern, D. Eck, J. Dean, S. Petrov, and N. Fiedel · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs, 2018
J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. VanderPlas, S. Wanderman-Milne, and Q. Zhang · 2018
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
J. Devlin, M. Chang, K. Lee, and K. Toutanova · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman · 2018
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Cited alongside, same era.
Scatterbrain: Unifying sparse and low-rank attention approximation
B. Chen, T. Dao, E. Winsor, Z. Song, A. Rudra, and C. Ré · 2021
Cited alongside, same era.
Rethinking attention with performers
K. M. Choromanski, V. Likhosherstov, D. Dohan, X. Song, A. Gane, T. Sarlos, P. Hawkins, J. Q. Davis, A. Mohiuddin, L. Kaiser, D. B. Belanger, L. J. Colwell, and A. Weller · 2021
Cited alongside, same era.
Later among the works it cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness, 2022
T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. Ré · 2022
Later among the works it cites.
Training compute-optimal large language models, 2022
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, J. W. Rae, O. Vinyals, and L. Sifre · 2022
Later among the works it cites.
Transformer quality in linear time, 2022
W. Hua, Z. Dai, H. Liu, and Q. V. Le · 2022
Later among the works it cites.
Linear complexity randomized self-attention mechanism
L. Zheng, C. Wang, and L. Kong · 2022
Later among the works it cites.
Introducing 100k context windows, 2023
Anthropic · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Efficient attention via control variates
L. Zheng, J. Yuan, C. Wang, and L. Kong · 2023
Closest in time.