Fetching the paper…
Reading the bibliography…
Over the past 7 years, attention has become one of the most important primitives in deep learning.
Exploring the limits of transfer learning with a unified text-to-text transformer, 2023
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 1910
Earlier work this paper cites.
Longformer: The long-document transformer, 2020a
Beltagy, I., Peters, M. E., and Cohan, A · 2004
Earlier work this paper cites.
Longformer: The long-document transformer, 2020b
Beltagy, I., Peters, M. E., and Cohan, A · 2004
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
Tvm: an automated end-to-end optimizing compiler for deep learning
Chen, T., Moreau, T., Jiang, Z., Zheng, L., Yan, E., Cowan, M., Shen, H., Wang, L., Hu, Y., Ceze, L., Guestrin, C., and Krishnamurthy, A · 2018
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with IO-awareness
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C · 2022
Earlier work this paper cites.
Dilated neighborhood attention transformer, 2022
Hassani, A. and Shi, H · 2022
Earlier work this paper cites.
Train short, test long: Attention with linear biases enables input length extrapolation, 2022
Press, O., Smith, N. A., and Lewis, M · 2022
Earlier work this paper cites.
Self-attention does not need o ( n 2 ) o(n^{2}) memory, 2022
Rabe, M. N. and Staats, C · 2022
Cited alongside, same era.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints, 2023
Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebrón, F., and Sanghai, S · 2023
Cited alongside, same era.
Accelerating generative ai with pytorch ii: Gpt, fast, November 2023
gpt-fast maintainers and contributors · 2023
Cited alongside, same era.
Neighborhood attention transformer
Hassani, A., Walton, S., Li, J., Li, S., and Shi, H · 2023
Cited alongside, same era.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Cited alongside, same era.
The llama 3 herd of models, 2024
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., and Akhil Mathur, e. a · 2024
Closest in time.
Hassani, A., Hwu, W.-M., and Shi, H · 2024
Closest in time.
Flashattention-3: Fast and accurate attention with asynchrony and low-precision, 2024
Shah, J., Bikshandi, G., Zhang, Y., Thakkar, V., Ramani, P., and Dao, T · 2024
Closest in time.
Gemma 2: Improving open language models at a practical size, 2024
Team, G., Riviere, M., Pathak, S., Sessa, P. G., Hardin, C., Bhupatiraju, S., Hussenot, L., and et al., T. M · 2024
Closest in time.
torchtune: Pytorch’s finetuning library, April 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pytorch 2: Faster machine learning through dynamic python bytecode transformation and graph compilation
Ansel, J., Yang, E., He, H., Gimelshein, N., Jain, A., Voznesensky, M., Bao, B., Bell, P., and et al · 2024
Cited alongside, same era.
Flashattention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2024
Cited alongside, same era.
Flashdecoding for long-context inference, 2023
Dao, T., Haziza, D., Massa, F., and Sizov, G · 2024
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., and Stoica, I
Cited in the paper.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., and Stoica, I
Cited in the paper.
torchtune maintainers and contributors · 2024
Closest in time.
Flashmask: Efficient and rich mask extension of flashattention
Wang, G., Zeng, J., Xiao, X., Wu, S., Yang, J., Zheng, L., Chen, Z., Bian, J., Yu, D., and Wang, H · 2024
Closest in time.
A multi-level superoptimizer for tensor programs, 2024
Wu, M., Cheng, X., Padon, O., and Jia, Z · 2024
Closest in time.