Fetching the paper…
Reading the bibliography…
Transformer-based large language models (LLMs) cache context as key-value (KV) pairs during inference.
Zip file format specification, 1989
P. W. Katz · 1989
Earlier work this paper cites.
Professor forcing: A new algorithm for training recurrent networks
A. Goyal, A. Lamb, Y. Zhang, S. Zhang, A. Courville, and Y. Bengio · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Deep contextualized word representations
M. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training, 2018
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al · 2018
Earlier work this paper cites.
Generating long sequences with sparse transformers
R. Child, S. Gray, A. Radford, and I. Sutskever · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Earlier work this paper cites.
Compressive transformers for long-range sequence modelling
J. W. Rae, A. Potapenko, S. M. Jayakumar, and T. P. Lillicrap · 2020
Earlier work this paper cites.
Big bird: Transformers for longer sequences
M. Zaheer, G. Guruganesh, K. A. Dubey, J. Ainslie, C. Alberti, S. Ontanon, P. Pham, A. Ravula, Q. Wang, L. Yang, et al · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Earlier work this paper cites.
Masked autoencoders are scalable vision learners
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick · 2022
Earlier work this paper cites.
Learned token pruning for transformers
S. Kim, S. Shen, D. Thorsley, A. Gholami, W. Kwon, J. Hassoun, and K. Keutzer · 2022
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Earlier work this paper cites.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints
J. Ainslie, J. Lee-Thorp, M. De Jong, Y. Zemlyanskiy, F. Lebrón, and S. Sanghai · 2023
Earlier work this paper cites.
Dynamic context pruning for efficient and interpretable autoregressive transformers
S. Anagnostidis, D. Pavllo, L. Biggio, L. Noci, A. Lucchi, and T. Hofmann · 2023
Earlier work this paper cites.
Mistral 7b, 2023
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, et al · 2023
Cited alongside, same era.
Needle in a haystack-pressure testing llms, 2023
G. Kamradt · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica · 2023
Cited alongside, same era.
Decodingtrust: A comprehensive assessment of trustworthiness in gpt models
B. Wang, W. Chen, H. Pei, C. Xie, M. Kang, et al · 2023
Cited alongside, same era.
H2o: Heavy-hitter oracle for efficient generative inference of large language models
Z. Zhang, Y. Sheng, T. Zhou, T. Chen, L. Zheng, et al · 2023
Cited alongside, same era.
Taming throughput-latency tradeoff in llm inference with sarathi-serve
A. Agrawal, N. Kedia, A. Panwar, J. Mohan, N. Kwatra, B. Gulavani, A. Tumanov, and R. Ramjee · 2024
Ruler: What’s the real context size of your long-context language models?
C.-P. Hsieh, S. Sun, S. Kriman, S. Acharya, D. Rekesh, F. Jia, Y. Zhang, and B. Ginsburg · 2024
Later among the works it cites.
Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention
H. Jiang, Y. Li, C. Zhang, Q. Wu, X. Luo, et al · 2024
Later among the works it cites.
Infinigen: Efficient generative inference of large language models with dynamic kv cache management
W. Lee, J. Lee, J. Seo, and J. Sim · 2024
Later among the works it cites.
Qserve: W4a8kv4 quantization and system co-design for efficient llm serving
Y. Lin, H. Tang, S. Yang, Z. Zhang, G. Xiao, C. Gan, and S. Han · 2024
Later among the works it cites.
Transformers are multi-state rnns
M. Oren, M. Hassid, N. Yarden, Y. Adi, and R. Schwartz · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Longbench: A bilingual, multitask benchmark for long context understanding
Y. Bai, X. Lv, J. Zhang, H. Lyu, J. Tang, et al · 2024
Cited alongside, same era.
Pyramidkv: Dynamic kv cache compression based on pyramidal information funneling
Z. Cai, Y. Zhang, B. Gao, Y. Liu, T. Liu, K. Lu, et al · 2024
Cited alongside, same era.
Don’t do rag: When cache-augmented generation is all you need for knowledge tasks
B. J. Chan, C.-T. Chen, J.-H. Cheng, and H.-H. Huang · 2024
Cited alongside, same era.
Optimizing ai inference at character.ai, 2024
Character.AI · 2024
Cited alongside, same era.
Nacl: A general and effective kv cache eviction framework for llms at inference time
Y. Chen, G. Wang, J. Shang, S. Cui, Z. Zhang, T. Liu, S. Wang, Y. Sun, D. Yu, and H. Wu · 2024
Cited alongside, same era.
Finch: Prompt-guided key-value cache compression for large language models
G. Corallo and P. Papotti · 2024
Cited alongside, same era.
Quest: Query-aware sparsity for efficient long-context llm inference
J. Tang, Y. Zhao, K. Zhu, G. Xiao, B. Kasikci, and S. Han · 2024
Later among the works it cites.
Efficient streaming language models with attention sinks
G. Xiao, Y. Tian, B. Chen, S. Han, and M. Lewis · 2024
Later among the works it cites.
∞ \infty bench: Extending long context evaluation beyond 100k tokens
X. Zhang, Y. Chen, S. Hu, Z. Xu, J. Chen, et al · 2024
Later among the works it cites.
Magicpig: Lsh sampling for efficient llm generation
Z. Chen, R. Sadhukhan, Z. Ye, Y. Zhou, J. Zhang, et al · 2025
Closest in time.
Beyond rag: Task-aware kv cache compression for comprehensive knowledge reasoning
G. Corallo, O. Weller, F. Petroni, and P. Papotti · 2025
Closest in time.
Scbench: A kv cache-centric analysis of long-context methods
Y. Li, H. Jiang, Q. Wu, X. Luo, S. Ahn, C. Zhang, A. H. Abdi, D. Li, J. Gao, Y. Yang, et al · 2025
Closest in time.
Safety alignment should be made more than just a few tokens deep
X. Qi, A. Panda, K. Lyu, X. Ma, S. Roy, A. Beirami, P. Mittal, and P. Henderson · 2025
Closest in time.
G. Team, A. Kamath, J. Ferret, S. Pathak, N. Vieillard, et al · 2025
Closest in time.
Duoattention: Efficient long-context llm inference with retrieval and streaming heads
G. Xiao, J. Tang, J. Zuo, J. Guo, S. Yang, H. Tang, Y. Fu, and S. Han · 2025
Closest in time.
A. Yang, B. Yu, C. Li, D. Liu, F. Huang, H. Huang, et al · 2025
Closest in time.
Differential transformer
T. Ye, L. Dong, Y. Xia, Y. Sun, Y. Zhu, G. Huang, and F. Wei · 2025
Closest in time.