Fetching the paper…
Reading the bibliography…
Effective attention modules have played a crucial role in the success of Transformer-based large language models (LLMs), but the quadratic time and memory complexities of these attention modules also pose a challenge when processing long sequences.
A bridging model for parallel computation
Valiant, L. G · 1990
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Chen, T., Xu, B., Zhang, C., and Guestrin, C · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Online normalizer calculation for softmax
Milakov, M. and Gimelshein, N · 2018
Earlier work this paper cites.
GPipe: efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, M. X., Chen, D., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., et al · 2019
Earlier work this paper cites.
Set Transformer: A framework for attention-based permutation-invariant neural networks
Lee, J., Lee, Y., Kim, J., Kosiorek, A., Choi, S., and Teh, Y. W · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Rethinking attention with performers
Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L., et al · 2020
Earlier work this paper cites.
Transformers are RNNs: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
ZeRO: Memory optimizations toward training trillion parameter models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y · 2020
Cited alongside, same era.
Linformer: Self-attention with linear complexity
Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H · 2020
Cited alongside, same era.
Lightweight and efficient end-to-end speech recognition using low-rank transformer
Winata, G. I., Cahyawijaya, S., Lin, Z., Liu, Z., and Fung, P · 2020
Cited alongside, same era.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Cited alongside, same era.
Pre-trained models: Past, present and future
Han, X., Zhang, Z., Ding, N., Gu, Y., Liu, X., Huo, Y., Qiu, J., Yao, Y., Zhang, A., Zhang, L., et al · 2021
Cited alongside, same era.
PaLM: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Later among the works it cites.
FlashAttention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D., Ermon, S., Rudra, A., and Ré, C · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Later among the works it cites.
The devil in linear transformer
Qin, Z., Han, X., Sun, W., Li, D., Kong, L., Barnes, N., and Zhong, Y · 2022
Later among the works it cites.
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Perceiver: General perception with iterative attention
Jaegle, A., Gimeno, F., Brock, A., Vinyals, O., Zisserman, A., and Carreira, J · 2021
Cited alongside, same era.
Sequence parallelism: Long sequence training from system perspective
Li, S., Xue, F., Baranwal, C., Li, Y., and You, Y · 2021
Cited alongside, same era.
Efficient large-scale language model training on gpu clusters using Megatron-LM
Narayanan, D., Shoeybi, M., Casper, J., LeGresley, P., Patwary, M., Korthikanti, V., Vainbrand, D., Kashinkunti, P., Bernauer, J., Catanzaro, B., et al · 2021
Cited alongside, same era.
Self-attention does not need o ( n 2 ) o(n^{2}) memory
Rabe, M. N. and Staats, C · 2021
Cited alongside, same era.
ZeRO-Offload: Democratizing billion-scale model training
Ren, J., Rajbhandari, S., Aminabadi, R. Y., Ruwase, O., Yang, S., Zhang, M., Li, D., and He, Y · 2021
Cited alongside, same era.
LLaMA: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al
Cited in the paper.
LLaMA 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al
Cited in the paper.
Flashattention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2023
Later among the works it cites.
LongNet: Scaling transformers to 1,000,000,000 tokens
Ding, J., Ma, S., Dong, L., Zhang, X., Huang, S., Wang, W., and Wei, F · 2023
Later among the works it cites.
Reducing activation recomputation in large transformer models
Korthikanti, V. A., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B · 2023
Later among the works it cites.
A survey of large language models
Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., et al · 2023
Later among the works it cites.