Fetching the paper…
Reading the bibliography…
The quadratic time and memory complexity inherent to self-attention mechanisms, with respect to sequence length, presents a critical computational bottleneck in the training and deployment of large-scale Transformer-based language models.
Compressive transformers for long-range sequence modelling
Rae, J. W., Potapenko, A., Jayakumar, S. M., Hillier, C., and Lillicrap, T. P · 1911
Earlier work this paper cites.
Prefix sums and their applications
Blelloch, G. E · 1990
Earlier work this paper cites.
Which problems have strongly exponential complexity?
Impagliazzo, R., Paturi, R., and Zane, F · 2001
Earlier work this paper cites.
Subspace embeddings for the polynomial kernel
Avron, H., Nguyen, H., and Woodruff, D · 2014
Earlier work this paper cites.
Sketching as a tool for numerical linear algebra
Woodruff, D. P. et al · 2014
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Language modeling with gated convolutional networks
Dauphin, Y. N., Fan, A., Auli, M., and Grangier, D · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
(learned) frequency estimation algorithms under zipfian distribution
Aamand, A., Indyk, P., and Vakilian, A · 2019
Earlier work this paper cites.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Learning-based frequency estimation algorithms
Hsu, C.-Y., Indyk, P., Katabi, D., and Vakilian, A · 2019
Earlier work this paper cites.
Tight dimensionality reduction for sketching low degree polynomial kernels
Meister, M., Sarlos, T., and Woodruff, D · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Transformer dissection: a unified understanding of transformer’s attention via the lens of kernel
Tsai, Y.-H. H., Bai, S., Yamada, M., Morency, L.-P., and Salakhutdinov, R · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Oblivious sketching of high-degree polynomial kernels
Ahle, T. D., Kapralov, M., Knudsen, J. B., Pagh, R., Velingker, A., Woodruff, D. P., and Zandieh, A · 2020
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Earlier work this paper cites.
PIQA: reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Bras, R. L., Gao, J., and Choi, Y · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Rethinking attention with performers
Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L., et al · 2020
Cited alongside, same era.
Wiki-40b: Multilingual language model dataset
Guo, M., Dai, Z., Vrandečić, D., and Al-Rfou, R · 2020
Cited alongside, same era.
Polynomial tensor sketch for element-wise function of low-rank matrix
Han, I., Avron, H., and Shin, J · 2020
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Later among the works it cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D., Ermon, S., Rudra, A., and Ré, C · 2022
Later among the works it cites.
Transformer quality in linear time
Hua, W., Dai, Z., Liu, H., and Le, Q · 2022
Later among the works it cites.
In-context learning and induction heads
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., et al · 2022
Later among the works it cites.
Efficient transformers: A survey
Tay, Y., Dehghani, M., Bahri, D., and Metzler, D · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Cited alongside, same era.
Reformer: The efficient transformer
Kitaev, N., Kaiser, Ł., and Levskaya, A · 2020
Cited alongside, same era.
Glu variants improve transformer
Shazeer, N · 2020
Cited alongside, same era.
Linformer: Self-attention with linear complexity
Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H · 2020
Cited alongside, same era.
Big bird: Transformers for longer sequences
Zaheer, M., Guruganesh, G., Dubey, K. A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., et al · 2020
Cited alongside, same era.
Finetuning pretrained transformers into rnns
Kasai, J., Peng, H., Zhang, Y., Yogatama, D., Ilharco, G., Pappas, N., Mao, Y., Chen, W., and Smith, N. A · 2021
Cited alongside, same era.
Peng, H., Pappas, N., Yogatama, D., Schwartz, R., Smith, N. A., and Kong, L · 2021
Cited alongside, same era.
Later among the works it cites.
Fast attention requires bounded entries
Alman, J. and Song, Z · 2023
Closest in time.
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al · 2023
Closest in time.
Linear complexity self-attention with 3rd order polynomials
Babiloni, F., Marras, I., Deng, J., Kokkinos, F., Maggioni, M., Chrysos, G., Torr, P., and Zafeiriou, S · 2023
Closest in time.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2023
Closest in time.
Flashattention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2023
Closest in time.
Longnet: Scaling transformers to 1,000,000,000 tokens
Ding, J., Ma, S., Dong, L., Zhang, X., Huang, S., Wang, W., Zheng, N., and Wei, F · 2023
Closest in time.
Mamba: Linear-time sequence modeling with selective state spaces
Gu, A. and Dao, T · 2023
Closest in time.
Hyperattention: Long-context attention in near-linear time
Han, I., Jayaram, R., Karbasi, A., Mirrokni, V., Woodruff, D. P., and Zandieh, A · 2023
Closest in time.
Implementation of FlashAttention in Pallas
JAX authors · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Efficient streaming language models with attention sinks
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2023
Closest in time.
Gated linear attention transformers with hardware-efficient training
Yang, S., Wang, B., Shen, Y., Panda, R., and Kim, Y · 2023
Closest in time.