Fetching the paper…
Reading the bibliography…
We present context parallelism for long-context large language model inference, which achieves near-linear scaling for long-context prefill latency with up to 128 H100 GPUs across 16 nodes.
Attention is all you need
Vaswani, A · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J · 2018
Earlier work this paper cites.
Online normalizer calculation for softmax
Milakov, M. and Gimelshein, N · 2018
Earlier work this paper cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, D., Chen, M., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., et al · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Fast transformer decoding: One write-head is all you need
Shazeer, N · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 2019
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H · 2020
Earlier work this paper cites.
Sequence parallelism: Long sequence training from system perspective
Li, S., Xue, F., Baranwal, C., Li, Y., and You, Y · 2021
Earlier work this paper cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B · 2021
Earlier work this paper cites.
Efficient large-scale language model training on gpu clusters using megatron-lm
Narayanan, D., Shoeybi, M., Casper, J., LeGresley, P., Patwary, M., Korthikanti, V., Vainbrand, D., Kashinkunti, P., Bernauer, J., Catanzaro, B., et al · 2021
Earlier work this paper cites.
Simple local attentions remain competitive for long-context tasks
Xiong, W., Oğuz, B., Gupta, A., Chen, X., Liskovich, D., Levy, O., Yih, W.-t., and Mehdad, Y · 2021
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D., Ermon, S., Rudra, A., and Ré, C · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
GQA: Training generalized multi-query transformer models from multi-head checkpoints
Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebrón, F., and Sanghai, S · 2023
Cited alongside, same era.
Striped attention: Faster ring attention for causal transformers
Brandon, W., Nrusimha, A., Qian, K., Ankner, Z., Jin, T., Song, Z., and Ragan-Kelley, J · 2023
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2023
Cited alongside, same era.
Flashattention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2023
Cited alongside, same era.
https://platform.openai.com/docs/models
GPT-4o with 128K context length · 2024
Closest in time.
Kv-runahead: Scalable causal llm inference by parallel key-value cache generation
Cho, M., Rastegari, M., and Naik, D · 2024
Closest in time.
Usp: A unified sequence parallelism approach for long context generative ai
Fang, J. and Zhao, S · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team · 2024
Closest in time.
Kvquant: Towards 10 million context length llm inference with kv cache quantization
Hooper, C., Kim, S., Mohammadzadeh, H., Mahoney, M. W., Shao, Y. S., Keutzer, K., and Gholami, A · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gemini Team · 2023
Cited alongside, same era.
Jacobs, S. A., Tanaka, M., Zhang, C., Zhang, M., Song, S. L., Rajbhandari, S., and He, Y · 2023
Cited alongside, same era.
Reducing activation recomputation in large transformer models
Korthikanti, V. A., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Cited alongside, same era.
Distflashattn: Distributed memory-efficient attention for long-context llms training
Li, D., Shao, R., Xie, A., Xing, E. P., Ma, X., Stoica, I., Gonzalez, J. E., and Zhang, H · 2023
Cited alongside, same era.
Ring attention with blockwise transformers for near-infinite context
Liu, H., Zaharia, M., and Abbeel, P · 2023
Cited alongside, same era.
Efficiently scaling transformer inference
Pope, R., Douglas, S., Chowdhery, A., Devlin, J., Bradbury, J., Heek, J., Xiao, K., Agrawal, S., and Dean, J · 2023
Cited alongside, same era.
Efficient streaming language models with attention sinks
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2023
Cited alongside, same era.
Closest in time.
Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention
Jiang, H., Li, Y., Zhang, C., Wu, Q., Luo, X., Ahn, S., Han, Z., Abdi, A. H., Li, D., Lin, C.-Y., et al · 2024
Closest in time.
Hydragen: High-throughput llm inference with shared prefixes
Juravsky, J., Brown, B., Ehrlich, R., Fu, D. Y., Ré, C., and Mirhoseini, A · 2024
Closest in time.
Qserve: W4a8kv4 quantization and system co-design for efficient llm serving
Lin, Y., Tang, H., Yang, S., Zhang, Z., Xiao, G., Gan, C., and Han, S · 2024
Closest in time.
Llama Team · 2024
Closest in time.
OCP summit 2022: Open hardware for ai infrastructure
Meta Engineering · 2024
Closest in time.
Leave no context behind: Efficient infinite context transformers with infini-attention
Munkhdalai, T., Faruqui, M., and Gopal, S · 2024
Closest in time.
Getting started with CUDA Graphs
Nvidia Blog · 2024
Closest in time.
Mooncake: A kvcache-centric disaggregated architecture for llm serving
Qin, R., Li, Z., He, W., Zhang, M., Wu, Y., Zheng, W., and Xu, X · 2024
Closest in time.
Flashattention-3: Fast and accurate attention with asynchrony and low-precision
Shah, J., Bikshandi, G., Zhang, Y., Thakkar, V., Ramani, P., and Dao, T · 2024
Closest in time.
Loongserve: Efficiently serving long-context large language models with elastic sequence parallelism
Wu, B., Liu, S., Zhong, Y., Sun, P., Liu, X., and Jin, X · 2024
Closest in time.
Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving, 2024
Zhong, Y., Liu, S., Chen, J., Hu, J., Zhu, Y., Liu, X., Jin, X., and Zhang, H · 2024
Closest in time.