Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) exhibit pronounced memory-bound characteristics during inference due to High Bandwidth Memory (HBM) bandwidth constraints.
Fast transformer decoding: One write-head is all you need
Shazeer, N. 2019 · 1911
Earlier work this paper cites.
Attention is All you Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L. u.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning
Rajbhandari, S.; Ruwase, O.; Rasley, J.; Smith, S.; and He, Y. 2021 · 2021
Earlier work this paper cites.
Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale
Aminabadi, R. Y.; Rajbhandari, S.; Awan, A. A.; Li, C.; Li, D.; Zheng, E.; Ruwase, O.; Smith, S.; Zhang, M.; Rasley, J.; et al. 2022 · 2022
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T.; Fu, D.; Ermon, S.; Rudra, A.; and Ré, C. 2022 · 2022
Earlier work this paper cites.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Earlier work this paper cites.
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Ainslie, J.; Lee-Thorp, J.; de Jong, M.; Zemlyanskiy, Y.; Lebr’on, F.; and Sanghai, S. K. 2023 · 2023
Earlier work this paper cites.
SeamlessM4T: Massively Multilingual & Multimodal Machine Translation
Barrault, L.; Chung, Y.-A.; Meglioli, M. C.; Dale, D.; Dong, N.; Duquenne, P.-A.; Elsahar, H.; Gong, H.; Heffernan, K.; Hoffman, J.; et al. 2023 · 2023
Cited alongside, same era.
Flashattention-2: Faster attention with better parallelism and work partitioning
Dao, T. 2023 · 2023
Cited alongside, same era.
Large language models as zero-shot conversational recommenders
He, Z.; Xie, Z.; Jha, R.; Steck, H.; Liang, D.; Feng, Y.; Majumder, B. P.; Kallus, N.; and McAuley, J. 2023 · 2023
Cited alongside, same era.
Deepum: Tensor migration and prefetching in unified memory
Jung, J.; Kim, J.; and Lee, J. 2023 · 2023
Cited alongside, same era.
Efficient Memory Management for Large Language Model Serving with PagedAttention
Kwon, W.; Li, Z.; Zhuang, S.; Sheng, Y.; Zheng, L.; Yu, C. H.; Gonzalez, J.; Zhang, H.; and Stoica, I. 2023 · 2023
Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model
Liu, A.; Feng, B.; Wang, B.; Wang, B.; Liu, B.; Zhao, C.; Dengr, C.; Ruan, C.; Dai, D.; Guo, D.; et al. 2024 · 2024
Later among the works it cites.
Splitwise: Efficient generative llm inference using phase splitting
Patel, P.; Choukse, E.; Zhang, C.; Shah, A.; Goiri, Í.; Maleki, S.; and Bianchini, R. 2024 · 2024
Later among the works it cites.
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
Shah, J.; Bikshandi, G.; Zhang, Y.; Thakkar, V.; Ramani, P.; and Dao, T. 2024 · 2024
Later among the works it cites.
CUDA Prefetch PTX Instruction
NVIDIA. 2025a · 2025
Closest in time.
PTX Cache Eviction Priority
NVIDIA. 2025b · 2025
Closest in time.
PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Fast Inference from Transformers via Speculative Decoding
Leviathan, Y.; Kalman, M.; and Matias, Y. 2023 · 2023
Cited alongside, same era.
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
Agrawal, A.; Kedia, N.; Panwar, A.; Mohan, J.; Kwatra, N.; Gulavani, B.; Tumanov, A.; and Ramjee, R. 2024 · 2024
Cited alongside, same era.
Yuzuguler, A. C.; Zhuang, J.; and Cavigelli, L. 2025 · 2025
Closest in time.