Fetching the paper…
Reading the bibliography…
Quantization has emerged as an effective and lightweight solution to reduce the memory footprint of the KV cache in Large Language Models.
Pointer sentinel mixture models, 2016
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
4 - automated methods
Sylvie Boldo and Guillaume Melquiond · 2017
Earlier work this paper cites.
Think you have solved question answering? Try ARC, the AI2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
MathQA: Towards interpretable math word problem solving with operation-based formalisms
Aida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi · 2019
Earlier work this paper cites.
PIQA: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding, 2021
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Efficiently scaling transformer inference, 2022
Reiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Anselm Levskaya, Jonathan Heek, Kefan Xiao, Shivani Agrawal, and Jeff Dean · 2022
Earlier work this paper cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al · 2023
Earlier work this paper cites.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al · 2023
Earlier work this paper cites.
Mistral 7b, 2023
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, et al · 2023
Earlier work this paper cites.
FlashAttention-2: Faster attention with better parallelism and work partitioning, 2023
Tri Dao · 2023
Earlier work this paper cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, et al · 2023
Earlier work this paper cites.
Roformer: Enhanced transformer with rotary position embedding, 2023
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu · 2023
Earlier work this paper cites.
GPTQ: Accurate post-training quantization for generative pre-trained transformers, 2023
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh · 2023
Earlier work this paper cites.
Flash-LLM: Enabling cost-effective and highly-efficient large generative model inference with unstructured sparsity, 2023
Haojun Xia, Zhen Zheng, Yuchao Li, Donglin Zhuang, Zhongzhu Zhou, Xiafei Qiu, Yong Li, Wei Lin, and Shuaiwen Leon Song · 2023
Earlier work this paper cites.
LongBench: A bilingual, multitask benchmark for long context understanding
Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li · 2024
Earlier work this paper cites.
RetrievalAttention: Accelerating long-context LLM inference via vector retrieval
Di Liu, Meng Chen, Baotong Lu, Huiqiang Jiang, Zhenhua Han, Qianxi Zhang, Qi Chen, Chengruidong Zhang, Bailu Ding, Kai Zhang, et al · 2024
Cited alongside, same era.
CaM: cache merging for memory-efficient LLMs inference
Yuxin Zhang, Yuxuan Du, Gen Luo, Yunshan Zhong, Zhenyu Zhang, Shiwei Liu, and Rongrong Ji · 2024
Cited alongside, same era.
KIVI: A tuning-free asymmetric 2bit quantization for KV cache
Zirui Liu, Jiayi Yuan, Hongye Jin, Shaochen Zhong, Zhaozhuo Xu, Vladimir Braverman, Beidi Chen, and Xia Hu · 2024
Cited alongside, same era.
SKVQ: Sliding-window key and value cache quantization for large language models, 2024
Haojie Duanmu, Zhihang Yuan, Xiuhong Li, Jiangfei Duan, Xingcheng Zhang, and Dahua Lin · 2024
Cited alongside, same era.
KVQuant: Towards 10 million context length LLM inference with KV cache quantization
Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W Mahoney, Sophia Shao, Kurt Keutzer, and Amir Gholami · 2024
Pruner-Zero: Evolving symbolic pruning metric from scratch for large language models, 2024
Peijie Dong, Lujun Li, Zhenheng Tang, Xiang Liu, Xinglin Pan, Qiang Wang, and Xiaowen Chu · 2024
Later among the works it cites.
Parzc: Parametric zero-cost proxies for efficient nas, 2024
Peijie Dong, Lujun Li, Xinglin Pan, Zimian Wei, Xiang Liu, Qiang Wang, and Xiaowen Chu · 2024
Later among the works it cites.
Atom: Low-bit quantization for efficient and accurate LLM serving, 2024
Yilong Zhao, Chien-Yu Lin, Kan Zhu, Zihao Ye, Lequn Chen, Size Zheng, Luis Ceze, Arvind Krishnamurthy, Tianqi Chen, and Baris Kasikci · 2024
Later among the works it cites.
Deepseek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
The lottery LLM hypothesis, rethinking what abilities should LLM compression preserve?
Zhenheng Tang, Xiang Liu, Qian Wang, Peijie Dong, Bingsheng He, Xiaowen Chu, and Bo Li · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Transformer-VQ: Linear-time transformers via vector quantization, 2024
Lucas D. Lingle · 2024
Cited alongside, same era.
KV cache is 1 bit per channel: Efficient large language model inference with coupled quantization
Tianyi Zhang, Jonah Yi, Zhaozhuo Xu, and Anshumali Shrivastava · 2024
Cited alongside, same era.
The Llama 3 herd of models, 2024
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, et al · 2024
Cited alongside, same era.
RoFormer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu · 2024
Cited alongside, same era.
Efficient streaming language models with attention sinks
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis · 2024
Cited alongside, same era.
PrefixQuant: Static quantization beats dynamic through prefixed outliers in LLMs
Mengzhao Chen, Yi Liu, Jiahao Wang, Yi Bin, Wenqi Shao, and Ping Luo · 2024
Cited alongside, same era.
SnapKV: LLM knows what you are looking for before generation, 2024
Yuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh, Acyr Locatelli, Hanchen Ye, Tianle Cai, Patrick Lewis, and Deming Chen · 2024
Cited alongside, same era.
Closest in time.
Can compressed llms truly act? An empirical evaluation of agentic capabilities in LLM compression
Peijie Dong, Zhenheng Tang, Xiang Liu, Lujun Li, Xiaowen Chu, and Bo Li · 2025
Closest in time.
MegaAgent: A large-scale autonomous LLM-based multi-agent system without predefined SOPs
Qian Wang, Tianyu Wang, Zhenheng Tang, Qinbin Li, Nuo Chen, Jingsheng Liang, and Bingsheng He · 2025
Closest in time.
Security and privacy challenges of large language models: A survey
Badhan Chandra Das, M Hadi Amini, and Yanzhao Wu · 2025
Closest in time.
AgentTaxo: Dissecting and benchmarking token distribution of LLM multi-agent systems
Qian Wang, Zhenheng Tang, Zichen Jiang, Nuo Chen, Tianyu Wang, and Bingsheng He · 2025
Closest in time.
OracleKV: Oracle guidance for question-independent KV cache compression
Yuanbing Zhu, Zhenheng Tang, Xiang Liu, Ang Li, Bo Li, Xiaowen Chu, and Bo Han · 2025
Closest in time.
A comprehensive survey on long context language modeling
Jiaheng Liu, Dawei Zhu, Zhiqi Bai, Yancheng He, Huanxuan Liao, Haoran Que, Zekun Wang, Chenchen Zhang, Ge Zhang, Jiebin Zhang, et al · 2025
Closest in time.
CacheBlend: Fast large language model serving for RAG with cached knowledge fusion
Jiayi Yao, Hanchen Li, Yuhan Liu, Siddhant Ray, Yihua Cheng, Qizheng Zhang, Kuntai Du, Shan Lu, and Junchen Jiang · 2025
Closest in time.
ChunkKV: Semantic-preserving kv cache compression for efficient long-context LLM inference, 2025
Xiang Liu, Zhenheng Tang, Peijie Dong, Zeyu Li, Bo Li, Xuming Hu, and Xiaowen Chu · 2025
Closest in time.
Can LLMs maintain fundamental abilities under KV cache compression?, 2025
Xiang Liu, Zhenheng Tang, Hong Chen, Peijie Dong, Zeyu Li, Xiuze Zhou, Bo Li, Xuming Hu, and Xiaowen Chu · 2025
Closest in time.
SampleAttention: Near-lossless acceleration of long context LLM inference with adaptive structured sparse attention
Qianchao Zhu, Jiangfei Duan, Chang Chen, Siran Liu, Xiuhong Li, Guanyu Feng, Xin Lv, Huanqi Cao, Chuanfu Xiao, Xingcheng Zhang, Dahua Lin, and Chao Yang · 2025
Closest in time.
QServe: W4A8KV4 quantization and system co-design for efficient LLM serving, 2025
Yujun Lin, Haotian Tang, Shang Yang, Zhekai Zhang, Guangxuan Xiao, Chuang Gan, and Song Han · 2025
Closest in time.
Spinfer: Leveraging low-level sparsity for efficient large language model inference on gpus
Ruibo Fan, Xiangrui Yu, Peijie Dong, Zeyu Li, Gu Gong, Qiang Wang, Wei Wang, and Xiaowen Chu · 2025
Closest in time.