Fetching the paper…
Reading the bibliography…
Key-Value cache (\texttt{KV} \texttt{cache}) compression has emerged as a promising technique to optimize Large Language Model (LLM) serving.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M · 2020
Earlier work this paper cites.
Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale
Aminabadi, R. Y., Rajbhandari, S., Awan, A. A., Li, C., Li, D., Zheng, E., Ruwase, O., Smith, S., Zhang, M., Rasley, J., et al · 2022
Earlier work this paper cites.
FlashAttention: Fast and memory-efficient exact attention with IO-awareness
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C · 2022
Earlier work this paper cites.
Longbench: A bilingual, multitask benchmark for long context understanding
Bai, Y., Lv, X., Zhang, J., Lyu, H., Tang, J., Huang, Z., Du, Z., Liu, X., Zeng, A., Hou, L., et al · 2023
Earlier work this paper cites.
Lmdeploy: A toolkit for compressing, deploying, and serving llm
Contributors, L · 2023
Earlier work this paper cites.
Model tells you what to discard: Adaptive kv cache compression for llms
Ge, S., Zhang, Y., Liu, L., Zhang, M., Han, J., and Gao, J · 2023
Earlier work this paper cites.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Earlier work this paper cites.
OpenAI · 2023
Earlier work this paper cites.
Flexgen: High-throughput generative inference of large language models with a single gpu
Sheng, Y., Zheng, L., Yuan, B., Li, Z., Ryabinin, M., Chen, B., Liang, P., Ré, C., Stoica, I., and Zhang, C · 2023
Earlier work this paper cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Earlier work this paper cites.
Efficient streaming language models with attention sinks
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2023
Earlier work this paper cites.
Response length perception and sequence scheduling: An llm-empowered llm inference pipeline
Zheng, Z., Ren, X., Xue, F., Luo, Y., Jiang, X., and You, Y · 2023
Earlier work this paper cites.
Keyformer: Kv cache reduction through key tokens selection for efficient generative inference
Adnan, M., Arunkumar, A., Jain, G., Nair, P., Soloveychik, I., and Kamath, P · 2024
Earlier work this paper cites.
Vidur: A large-scale simulation framework for llm inference, 2024
Agrawal, A., Kedia, N., Mohan, J., Panwar, A., Kwatra, N., Gulavani, B., Ramjee, R., and Tumanov, A · 2024
Earlier work this paper cites.
Sharegpt vicuna unfiltered dataset
Anon · 2024
Earlier work this paper cites.
Claude ai, 2024
Anthropic · 2024
Cited alongside, same era.
Quarot: Outlier-free 4-bit inference in rotated llms
Ashkboos, S., Mohtashami, A., Croci, M. L., Li, B., Jaggi, M., Alistarh, D., Hoefler, T., and Hensman, J · 2024
Cited alongside, same era.
Palu: Compressing kv-cache with low-rank projection
Chang, C.-C., Lin, W.-C., Lin, C.-Y., Chen, C.-Y., Hu, Y.-F., Wang, P.-S., Huang, N.-C., Ceze, L., and Wu, K.-C · 2024
Cited alongside, same era.
Nacl: A general and effective kv cache eviction framework for llms at inference time
Chen, Y., Wang, G., Shang, J., Cui, S., Zhang, Z., Liu, T., Wang, S., Sun, Y., Yu, D., and Wu, H · 2024
Cited alongside, same era.
Sequence can secretly tell you what to discard
Dai, J., Huang, Z., Jiang, H., Chen, C., Cai, D., Bi, W., and Shi, S · 2024
Keep the cost down: A review on methods to optimize llm’s kv-cache consumption
Luohe, S., Hongyi, Z., Yao, Y., Zuchao, L., and Hai, Z · 2024
Later among the works it cites.
Transformers are multi-state rnns
Oren, M., Hassid, M., Adi, Y., and Schwartz, R · 2024
Later among the works it cites.
Efficient interactive llm serving with proxy model-based sequence length prediction
Qiu, H., Mao, W., Patke, A., Cui, S., Jha, S., Wang, C., Franke, H., Kalbarczyk, Z. T., Başar, T., and Iyer, R. K · 2024
Later among the works it cites.
On the efficacy of eviction policy for key-value constrained generative language model inference
Ren, S. and Zhu, K. Q · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
FlashAttention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2024
Cited alongside, same era.
Skvq: Sliding-window key and value cache quantization for large language models
Duanmu, H., Yuan, Z., Li, X., Duan, J., Zhang, X., and Lin, D · 2024
Cited alongside, same era.
Ada-kv: Optimizing kv cache eviction by adaptive budget allocation for efficient llm inference, 2024
Feng, Y., Lv, J., Cao, Y., Xie, X., and Zhou, S. K · 2024
Cited alongside, same era.
Flashinfer: A lightweight framework for inferencing
Flashinfer · 2024
Cited alongside, same era.
Lazyllm: Dynamic token pruning for efficient long context llm inference
Fu, Q., Cho, M., Merth, T., Mehta, S., Rastegari, M., and Najibi, M · 2024
Cited alongside, same era.
Gong, R., Yong, Y., Gu, S., Huang, Y., Zhang, Y., Liu, X., and Tao, D · 2024
Cited alongside, same era.
Introducing gemini: our largest and most capable ai model, 2023
Hassabis, D. and the Gemini Team · 2024
Cited alongside, same era.
Shi, Z., Ming, Y., Nguyen, X.-P., Liang, Y., and Joty, S · 2024
Later among the works it cites.
Squeezeattention: 2d management of kv-cache in llm inference via layer-wise optimal budget
Wang, Z. and Gan, S · 2024
Later among the works it cites.
Infllm: Training-free long-context extrapolation for llms with an efficient context memory
Xiao, C., Zhang, P., Han, X., Xiao, G., Lin, Y., Zhang, Z., Liu, Z., and Sun, M · 2024
Later among the works it cites.
Think: Thinner key cache by query-driven pruning
Xu, Y., Jie, Z., Dong, H., Wang, L., Lu, X., Zhou, A., Saha, A., Xiong, C., and Sahoo, D · 2024
Later among the works it cites.
Yuan, J., Liu, H., Chuang, Y.-N., Li, S., Wang, G., Le, D., Jin, H., Chaudhary, V., Xu, Z., Liu, Z., et al · 2024
Later among the works it cites.
Wkvquant: Quantizing weight and key/value cache for large language models gains more
Yue, Y., Yuan, Z., Duanmu, H., Zhou, S., Wu, J., and Nie, L · 2024
Later among the works it cites.
Qjl: 1-bit quantized jl transform for kv cache quantization with zero overhead
Zandieh, A., Daliri, M., and Han, I · 2024
Later among the works it cites.
Zero-delay qkv compression for mitigating kv cache and network bottlenecks in llm inference
Zhang, Z. and Shen, H · 2024
Later among the works it cites.
Zhu, Q., Duan, J., Chen, C., Liu, S., Li, X., Feng, G., Lv, X., Cao, H., Chuanfu, X., Zhang, X., Lin, D., and Yang, C · 2024
Later among the works it cites.
https://github.com/jy-yuan/KIVI/issues/4
Issue #4: [integrate kivi into inference frameworks?] · 2025
Closest in time.
Benchmarking llm inference backends
BentoML · 2025
Closest in time.