Fetching the paper…
Reading the bibliography…
Inference with Transformer-based Large Language Models (LLMs) on long sequences is both costly and slow due to the quadratic complexity of the self-attention mechanism.
Online normalizer calculation for softmax
Milakov, M. and Gimelshein, N · 2018
Earlier work this paper cites.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Earlier work this paper cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q., and Salakhutdinov, R · 2019
Earlier work this paper cites.
GPipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, D., Chen, M., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., et al · 2019
Earlier work this paper cites.
Megatron-LM: Training multi-billion parameter language models using model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 2019
Earlier work this paper cites.
Longformer: The long-document Transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., et al · 2020
Earlier work this paper cites.
Fully sharded data parallel: faster AI training with fewer GPUs, 2021
Meta-AI · 2021
Earlier work this paper cites.
FlashAttention: Fast and memory-efficient exact attention with IO-awareness
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C · 2022
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Earlier work this paper cites.
Sequence parallelism: Long sequence training from system perspective
Li, S., Xue, F., Baranwal, C., Li, Y., and You, Y · 2023
Earlier work this paper cites.
Blockwise parallel transformers for large context models
Liu, H. and Abbeel, P · 2023
Earlier work this paper cites.
Tensorrt-llm: An optimized library for large language models, 2023
NVIDIA · 2023
Cited alongside, same era.
H2o: Heavy-hitter oracle for efficient generative inference of large language models
Zhang, Z., Sheng, Y., Zhou, T., Chen, T., Zheng, L., Cai, R., Song, Z., Tian, Y., Re, C., Barrett, C., Wang, Z., and Chen, B · 2023
Cited alongside, same era.
The Claude 3 model family: Opus, Sonnet, Haiku, 2024
Anthropic · 2024
Cited alongside, same era.
Titans: Learning to memorize at test time
Behrouz, A., Zhong, P., and Mirrokni, V · 2024
Cited alongside, same era.
FlashAttention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2024
Cited alongside, same era.
Model tells you what to discard: Adaptive kv cache compression for llms
E2llm: Encoder elongated large language models for long-context understanding and reasoning
Liao, Z., Wang, J., Yu, H., Wei, L., Li, J., Wang, J., and Zhang, W · 2024
Closest in time.
Introducing Llama 3.1: Our most capable models to date, 2024
Meta-AI · 2024
Closest in time.
Leave no context behind: Efficient infinite context Transformers with infini-attention
Munkhdalai, T., Faruqui, M., and Gopal, S · 2024
Closest in time.
Lightning attention-2: A free lunch for handling unlimited sequence lengths in large language models
Qin, Z., Sun, W., Li, D., Shen, X., Sun, W., and Zhong, Y · 2024
Closest in time.
Writing in the margins: Better inference pattern for long context retrieval
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ge, S., Zhang, Y., Liu, L., Zhang, M., Han, J., and Gao, J · 2024
Cited alongside, same era.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini-Team · 2024
Cited alongside, same era.
RULER vs. Gradient’s 1M context length Llama-3-70B, 2024
Gradient.ai · 2024
Cited alongside, same era.
Lm-infinite: Zero-shot extreme length generalization for large language models
Han, C., Wang, Q., Peng, H., Xiong, W., Chen, Y., Ji, H., and Wang, S · 2024
Cited alongside, same era.
RULER: What’s the real context size of your long-context language models?
Hsieh, C.-P., Sun, S., Kriman, S., Acharya, S., Rekesh, D., Jia, F., and Ginsburg, B · 2024
Cited alongside, same era.
MInference 1.0: Accelerating pre-filling for long-context LLMs via dynamic sparse attention
Jiang, H., Li, Y., Zhang, C., Wu, Q., Luo, X., Ahn, S., Han, Z., Abdi, A. H., Li, D., Lin, C.-Y., Yang, Y., and Qiu, L · 2024
Cited alongside, same era.
BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack
Kuratov, Y., Bulatov, A., Anokhin, P., Rodkin, I., Sorokin, D., Sorokin, A., and Burtsev, M · 2024
Cited alongside, same era.
Russak, M., Jamil, U., Bryant, C., Kamble, K., Magnuson, A., Russak, M., and AlShikh, W · 2024
Closest in time.
Tree attention: Topology-aware decoding for long-context attention on gpu clusters
Shyam, V., Pilault, J., Shepperd, E., Anthony, Q., and Millidge, B · 2024
Closest in time.
You only cache once: Decoder-decoder architectures for language models
Sun, Y., Dong, L., Zhu, Y., Huang, S., Wang, W., Ma, S., Zhang, Q., Wang, J., and Wei, F · 2024
Closest in time.
Quest: query-aware sparsity for efficient long-context llm inference
Tang, J., Zhao, Y., Zhu, K., Xiao, G., Kasikci, B., and Han, S · 2024
Closest in time.
Scope: Optimizing key-value cache compression in long-context generation
Wu, J., Wang, Z., Zhang, L., Lai, Y., He, Y., and Zhou, D · 2024
Closest in time.
∞ \infty Bench: Extending long context evaluation beyond 100K tokens
Zhang, X., Chen, Y., Hu, S., Xu, Z., Chen, J., Hao, M., Han, X., Thai, Z., Wang, S., Liu, Z., and Sun, M · 2024
Closest in time.
Alisa: Accelerating large language model inference via sparsity-aware kv caching
Zhao, Y., Wu, D., and Wang, J · 2024
Closest in time.
Qwen · 2025
Closest in time.