Fetching the paper…
Reading the bibliography…
Advanced Large Language Models (LLMs) have achieved impressive performance across a wide range of complex and long-context natural language tasks.
Language models are few-shot learners. In Proceedings of the 34th International Conference on Neural Information Processing Systems . 1877–1901
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Linguistic regularities in sparse and explicit word representations. In Proceedings of the eighteenth conference on computational natural language learning . 171–180
Omer Levy and Yoav Goldberg. 2014 · 2014
Earlier work this paper cites.
Attention is all you need. In Advances in neural information processing systems . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Few-shot NLG with pre-trained language model
Zhiyu Chen, Harini Eavani, Wenhu Chen, Yinyin Liu, and William Yang Wang. 2019 · 2019
Earlier work this paper cites.
Compressive transformers for long-range sequence modelling
Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, and Timothy P Lillicrap. 2019 · 2019
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing. arXiv
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al · 2019
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan. 2020 · 2020
Earlier work this paper cites.
Longbench: A bilingual, multitask benchmark for long context understanding
Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, et al · 2023
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2023
Earlier work this paper cites.
Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs. In Workshop on Advancing Neural Network Training: Computational Efficiency, Scalability, and Resource Optimization (WANT@ NeurIPS 2023)
Suyu Ge, Yunan Zhang, Liyuan Liu, Minjia Zhang, Jiawei Han, and Jianfeng Gao. [n. d.] · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th Symposium on Operating Systems Principles . 611–626
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023 · 2023
Earlier work this paper cites.
Efficiently scaling transformer inference
Reiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Jonathan Heek, Kefan Xiao, Shivani Agrawal, and Jeff Dean. 2023 · 2023
Earlier work this paper cites.
Flexgen: High-throughput generative inference of large language models with a single gpu. In International Conference on Machine Learning . PMLR, 31094–31116
Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Ré, Ion Stoica, and Ce Zhang. 2023 · 2023
Earlier work this paper cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Earlier work this paper cites.
H2o: Heavy-hitter oracle for efficient generative inference of large language models
Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, et al · 2023
Earlier work this paper cites.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al · 2023
Earlier work this paper cites.
Keyformer: Kv cache reduction through key tokens selection for efficient generative inference
Muhammad Adnan, Akhil Arunkumar, Gaurav Jain, Prashant Nair, Ilya Soloveychik, and Purushotham Kamath. 2024 · 2024
Earlier work this paper cites.
Amey Agrawal, Junda Chen, Íñigo Goiri, Ramachandran Ramjee, Chaojie Zhang, Alexey Tumanov, and Esha Choukse. 2024 · 2024
Earlier work this paper cites.
Claude 3 Family Announcement
Anthropic. 2024 · 2024
Earlier work this paper cites.
Real-World Use Cases for Large Language Models (LLMs)
CellStrat. 2023 · 2024
Earlier work this paper cites.
7 Top Large Language Model Use Cases And Applications
Daivi. 2024 · 2024
Earlier work this paper cites.
QAQ: Quality Adaptive Quantization for LLM KV Cache
Shichen Dong, Wen Cheng, Jiayu Qin, and Wei Wang. 2024 · 2024
Cited alongside, same era.
{ \{ Cost-Efficient } \} large language model serving for multi-turn conversations with { \{ CachedAttention } \} . In 2024 USENIX Annual Technical Conference (USENIX ATC 24) . 111–126
Bin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang, Djordje Jevdjic, Junbo Deng, Xingkun Yang, Zhou Yu, and Pengfei Zuo. 2024 · 2024
Cited alongside, same era.
Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs. In The Twelfth International Conference on Learning Representations
Suyu Ge, Yunan Zhang, Liyuan Liu, Minjia Zhang, Jiawei Han, and Jianfeng Gao. 2024 · 2024
Cited alongside, same era.
Kvquant: Towards 10 million context length llm inference with kv cache quantization
Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami. 2024 · 2024
Cited alongside, same era.
Yarn-Llama-2-7B-128K-GGML
TheBloke. 2024 · 2024
Later among the works it cites.
So you want your private LLM at home?: a survey and benchmark of methods for efficient GPTs. In 11th IEEE Swiss Conference on Data Science (SDS), Zurich, Switzerland, 30-31 May 2024 . ZHAW Zürcher Hochschule für Angewandte Wissenschaften
Lukas Tuggener, Pascal Sager, Yassine Taoudi-Benchekroun, Benjamin F Grewe, and Thilo Stadelmann. 2024 · 2024
Later among the works it cites.
Look-m: Look-once optimization in kv cache for efficient multimodal long-context inference
Zhongwei Wan, Ziang Wu, Che Liu, Jinfa Huang, Zhihong Zhu, Peng Jin, Longyue Wang, and Li Yuan. 2024 · 2024
Later among the works it cites.
Leave No Document Behind: Benchmarking Long-Context LLMs with Extended Multi-Doc QA
Minzheng Wang, Longze Chen, Cheng Fu, Shengyi Liao, Xinghua Zhang, Bingli Wu, Haiyang Yu, Nan Xu, Lei Zhang, Run Luo, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mingqiang Huang, Ao Shen, Kai Li, Haoxiang Peng, Boyu Li, and Hao Yu. 2024a · 2024
Cited alongside, same era.
Locret: Enhancing Eviction in Long-Context LLM Inference with Trained Retaining Heads
Yuxiang Huang, Binhang Yuan, Xu Han, Chaojun Xiao, and Zhiyuan Liu. 2024b · 2024
Cited alongside, same era.
Hydragen: High-throughput llm inference with shared prefixes
Jordan Juravsky, Bradley Brown, Ryan Ehrlich, Daniel Y Fu, Christopher Ré, and Azalia Mirhoseini. 2024 · 2024
Cited alongside, same era.
12 Practical Large Language Model (LLM) Applications
Tim Keary. 2024 · 2024
Cited alongside, same era.
{ \{ InfiniGen } \} : Efficient generative inference of large language models with dynamic { \{ KV } \} cache management. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) . 155–172
Wonbeom Lee, Jungi Lee, Junghwan Seo, and Jaewoong Sim. 2024 · 2024
Cited alongside, same era.
Personal llm agents: Insights and survey about the capability, efficiency and security
Yuanchun Li, Hao Wen, Weijun Wang, Xiangyu Li, Yizhen Yuan, Guohong Liu, Jiacheng Liu, Wenxing Xu, Xiang Wang, Yi Sun, et al · 2024
Cited alongside, same era.
Accelerating Model Training in Multi-cluster Environments with Consumer-grade GPUs. In Proceedings of the ACM SIGCOMM 2024 Conference . 707–720
Hwijoon Lim, Juncheol Ye, Sangeetha Abdu Jyothi, and Dongsu Han. 2024 · 2024
Cited alongside, same era.
Infinite-llm: Efficient llm service for long context with distattention and distributed kvcache
Bin Lin, Tao Peng, Chen Zhang, Minmin Sun, Lanbo Li, Hanyu Zhao, Wencong Xiao, Qi Xu, Xiafei Qiu, Shen Li, et al · 2024
Cited alongside, same era.
Bingyang Wu, Shengyu Liu, Yinmin Zhong, Peng Sun, Xuanzhe Liu, and Xin Jin. 2024 · 2024
Later among the works it cites.
Think: Thinner key cache by query-driven pruning
Yuhui Xu, Zhanming Jie, Hanze Dong, Lei Wang, Xudong Lu, Aojun Zhou, Amrita Saha, Caiming Xiong, and Doyen Sahoo. 2024 · 2024
Later among the works it cites.
PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference
Dongjie Yang, XiaoDong Han, Yan Gao, Yao Hu, Shilin Zhang, and Hai Zhao. 2024 · 2024
Later among the works it cites.
ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 11608–11620
Lu Ye, Ze Tao, Yong Huang, and Yang Li. 2024 · 2024
Later among the works it cites.
Pqcache: Product quantization-based kvcache for long context llm inference
Hailin Zhang, Xiaodong Ji, Yilin Chen, Fangcheng Fu, Xupeng Miao, Xiaonan Nie, Weipeng Chen, and Bin Cui. 2024a · 2024
Later among the works it cites.
Enabling on-device llms personalization with smartphone sensing. In Companion of the 2024 on ACM International Joint Conference on Pervasive and Ubiquitous Computing . 186–190
Shiquan Zhang, Ying Ma, Le Fang, Hong Jia, Simon D’Alfonso, and Vassilis Kostakos. 2024c · 2024
Later among the works it cites.
Youpeng Zhao, Di Wu, and Jun Wang. 2024 · 2024
Later among the works it cites.
Sglang: Efficient execution of structured language model programs
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Livia Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E Gonzalez, et al · 2024
Later among the works it cites.
A survey on efficient inference for large language models
Zixuan Zhou, Xuefei Ning, Ke Hong, Tianyu Fu, Jiaming Xu, Shiyao Li, Yuming Lou, Luning Wang, Zhihang Yuan, Xiuhong Li, et al · 2024
Later among the works it cites.
Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
Qianchao Zhu, Jiangfei Duan, Chang Chen, Siran Liu, Xiuhong Li, Guanyu Feng, Xin Lv, Huanqi Cao, Xiao Chuanfu, Xingcheng Zhang, et al · 2024
Later among the works it cites.
{ \{ IMPRESS } \} : An { \{ Importance-Informed } \} { \{ Multi-Tier } \} Prefix { \{ KV } \} Storage System for Large Language Model Inference. In 23rd USENIX Conference on File and Storage Technologies (FAST 25) . 187–201
Weijian Chen, Shuibing He, Haoyang Qu, Ruidong Zhang, Siling Yang, Ping Chen, Yi Zheng, Baoxing Huai, and Gang Chen. 2025 · 2025
Closest in time.
LongChat: An Open Framework for Long-Context Language Models
Dacheng Li. 2025 · 2025
Closest in time.
LongChat-7B-v1.5-32k
LMSYS and Hugging Face. 2025 · 2025
Closest in time.
Yarn-Llama-2-13b-128k
NousResearch and Hugging Face. 2025 · 2025
Closest in time.
Xiurui Pan, Endian Li, Qiao Li, Shengwen Liang, Yizhou Shan, Ke Zhou, Yingwei Luo, Xiaolin Wang, and Jie Zhang. 2025 · 2025
Closest in time.
Jiayi Yao, Hanchen Li, Yuhan Liu, Siddhant Ray, Yihua Cheng, Qizheng Zhang, Kuntai Du, Shan Lu, and Junchen Jiang. 2025 · 2025
Closest in time.