Fetching the paper…
Reading the bibliography…
The deployment of efficient long-context LLMs in applications like autonomous agents, long-chain reasoning, and creative writing is fundamentally bottlenecked by the linear growth of KV cache memory.
Training deep nets with sublinear memory cost
Chen, T., Xu, B., Zhang, C., and Guestrin, C · 2016
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al · 2020
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K. R., and Cao, Y · 2022
Earlier work this paper cites.
Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models
Guha, N., Nyarko, J., Ho, D., Ré, C., Chilton, A., Chohlas-Wood, A., Peters, A., Waldon, B., Rockmore, D., Zambrano, D., et al · 2023
Earlier work this paper cites.
Swe-bench: Can language models resolve real-world github issues?
Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Earlier work this paper cites.
Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time
Liu, Z., Desai, A., Liao, F., Wang, W., Xie, V., Xu, Z., Kyrillidis, A., and Shrivastava, A · 2023
Earlier work this paper cites.
H2o: Heavy-hitter oracle for efficient generative inference of large language models
Zhang, Z., Sheng, Y., Zhou, T., Chen, T., Zheng, L., Cai, R., Song, Z., Tian, Y., Ré, C., Barrett, C., et al · 2023
Earlier work this paper cites.
Longbench: A bilingual, multitask benchmark for long context understanding
Bai, Y., Lv, X., Zhang, J., Lyu, H., Tang, J., Huang, Z., Du, Z., Liu, X., Zeng, A., Hou, L., et al · 2024
Earlier work this paper cites.
Pyramidkv: Dynamic kv cache compression based on pyramidal information funneling
Cai, Z., Zhang, Y., Gao, B., Liu, Y., Li, Y., Liu, T., Lu, K., Xiong, W., Dong, Y., Hu, J., et al · 2024
Earlier work this paper cites.
Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Vaughan, A., et al · 2024
Earlier work this paper cites.
Zipcache: Accurate and efficient kv cache quantization with salient token identification
He, Y., Zhang, L., Wu, W., Liu, J., Zhou, H., and Zhuang, B · 2024
Earlier work this paper cites.
Kvquant: Towards 10 million context length llm inference with kv cache quantization
Hooper, C., Kim, S., Mohammadzadeh, H., Mahoney, M. W., Shao, Y. S., Keutzer, K., and Gholami, A · 2024
Cited alongside, same era.
Lexico: Extreme kv cache compression via sparse coding over universal dictionaries
Kim, J., Park, J., Cho, J., and Papailiopoulos, D · 2024
Cited alongside, same era.
Cachegen: Kv cache compression and streaming for fast large language model serving
Liu, Y., Li, H., Cheng, Y., Ray, S., Huang, Y., Zhang, Q., Du, K., Yao, J., Lu, S., Ananthanarayanan, G., et al · 2024
Cited alongside, same era.
Transformers are multi-state rnns
Oren, M., Hassid, M., Yarden, N., Adi, Y., and Schwartz, R · 2024
Cited alongside, same era.
The fineweb datasets: Decanting the web for the finest text data at scale
Penedo, G., Kydlíček, H., Lozhkov, A., Mitchell, M., Raffel, C. A., Von Werra, L., Wolf, T., et al · 2024
Palu: Kv-cache compression with low-rank projection
Chang, C.-C., Lin, W.-C., Lin, C.-Y., Chen, C.-Y., Hu, Y.-F., Wang, P.-S., Huang, N.-C., Ceze, L., Abdelfattah, M. S., and Wu, K.-C · 2025
Later among the works it cites.
Expected attention: Kv cache compression by estimating attention from future queries distribution
Devoto, A., Jeblick, M., and Jégou, S · 2025
Later among the works it cites.
Ada-KV: Optimizing KV cache eviction by adaptive budget allocation for efficient LLM inference
Feng, Y., Lv, J., Cao, Y., Xie, X., and Zhou, S. K · 2025
Later among the works it cites.
Past-future scheduler for llm serving under sla guarantees
Gong, R., Bai, S., Wu, S., Fan, Y., Wang, Z., Li, X., Yang, H., and Liu, X · 2025
Later among the works it cites.
Efficient long-context llm inference via kv cache clustering
Hu, J., Wang, S., He, Y., Gong, P., Yi, J., Zhang, J., Bai, Y., Chen, R., Zhang, G., Li, C., et al · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Eigen attention: Attention in low-rank space for kv cache compression
Saxena, U., Saha, G., Choudhary, S., and Roy, K · 2024
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2024
Cited alongside, same era.
Shadowkv: Kv cache in shadows for high-throughput long-context llm inference
Sun, H., Chang, L.-W., Bao, W., Zheng, S., Zheng, N., Liu, X., Dong, H., Chi, Y., and Chen, B · 2024
Cited alongside, same era.
Quest: Query-aware sparsity for efficient long-context llm inference
Tang, J., Zhao, Y., Zhu, K., Xiao, G., Kasikci, B., and Han, S · 2024
Cited alongside, same era.
Team, Q. et al · 2024
Cited alongside, same era.
Infllm: Training-free long-context extrapolation for llms with an efficient context memory
Xiao, C., Zhang, P., Han, X., Xiao, G., Lin, Y., Zhang, Z., Liu, Z., and Sun, M · 2024
Cited alongside, same era.
Lorc: Low-rank compression for llms kv cache with a progressive compression strategy
Zhang, R., Wang, K., Liu, L., Wang, S., Cheng, H., Zhang, C., and Shen, Y · 2024
Cited alongside, same era.
Later among the works it cites.
Quantization meets dllms: A systematic study of post-training quantization for diffusion llms
Lin, H., Xu, H., Wu, Y., Guo, Z., Zhang, R., Lu, Z., Wei, Y., Zhang, Q., and Sun, Z · 2025
Later among the works it cites.
Clusterkv: Manipulating llm kv cache in semantic space for recallable compression
Liu, G., Li, C., Zhao, J., Zhang, C., and Guo, M · 2025
Later among the works it cites.
The sparse frontier: Sparse attention trade-offs in transformer llms
Nawrot, P., Li, R., Huang, R., Ruder, S., Marchisio, K., and Ponti, E. M · 2025
Later among the works it cites.
Deltallm: A training-free framework exploiting temporal sparsity for efficient edge llm inference
Qi, J., Gao, C., Ren, Z., and Chen, Q · 2025
Later among the works it cites.
Yang, A., Yu, B., Li, C., Liu, D., Huang, F., Huang, H., Jiang, J., Tu, J., Zhang, J., Zhou, J., et al · 2025
Later among the works it cites.
Pqcache: Product quantization-based kvcache for long context llm inference
Zhang, H., Ji, X., Chen, Y., Fu, F., Miao, X., Nie, X., Chen, W., and Cui, B · 2025
Later among the works it cites.
Maa invitational competitions
MAA · 2026
Closest in time.