Fetching the paper…
Reading the bibliography…
Deploying large language models (LLMs) for long-context inference remains challenging due to their substantial memory and computational demands.
Fast transformer decoding: One write-head is all you need
Noam Shazeer. 2019 · 1911
Earlier work this paper cites.
Who belongs in the family?
Robert L Thorndike. 1953 · 1953
Earlier work this paper cites.
The application of electronic computers to factor analysis
Henry F Kaiser. 1960 · 1960
Earlier work this paper cites.
On measures of entropy and information
Alfréd Rényi. 1961 · 1961
Earlier work this paper cites.
Principal components analysis (pca)
Andrzej Maćkiewicz and Waldemar Ratajczak. 1993 · 1993
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
Geoffrey E. Hinton, Simon Osindero, and Yee Whye Teh. 2006 · 2006
Earlier work this paper cites.
Scaling learning algorithms towards AI
Yoshua Bengio and Yann LeCun. 2007 · 2007
Earlier work this paper cites.
The effective rank: A measure of effective dimensionality
Olivier Roy and Martin Vetterli. 2007 · 2007
Earlier work this paper cites.
Pearson correlation coefficient
Israel Cohen, Yiteng Huang, Jingdong Chen, Jacob Benesty, Jacob Benesty, Jingdong Chen, Yiteng Huang, and Israel Cohen. 2009 · 2009
Earlier work this paper cites.
Mathematische grundlagen der quantenmechanik , volume 38
John Von Neumann. 2013 · 2013
Earlier work this paper cites.
Measures of entropy from data using infinitely divisible kernels
Luis Gonzalo Sanchez Giraldo, Murali Rao, and Jose C Principe. 2014 · 2014
Earlier work this paper cites.
Deep learning , volume 1
Ian Goodfellow, Yoshua Bengio, Aaron Courville, and Yoshua Bengio. 2016 · 2016
Earlier work this paper cites.
The wikitext long term dependency language modeling dataset
Stephen Merity. 2016 · 2016
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021 · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021 · 2021
Earlier work this paper cites.
Rank diminishing in deep neural networks
Ruili Feng, Kecheng Zheng, Yukun Huang, Deli Zhao, Michael Jordan, and Zheng-Jun Zha. 2022 · 2022
Earlier work this paper cites.
Language model compression with weighted low-rank factorization
Yen-Chang Hsu, Ting Hua, Sungen Chang, Qian Lou, Yilin Shen, and Hongxia Jin. 2022 · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Earlier work this paper cites.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai. 2023 · 2023
Cited alongside, same era.
Longbench: A bilingual, multitask benchmark for long context understanding
Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, et al. 2023 · 2023
Cited alongside, same era.
Flashattention-2: Faster attention with better parallelism and work partitioning
Tri Dao. 2023 · 2023
Cited alongside, same era.
Language modeling is compression
Grégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt, Tim Genewein, Christopher Mattern, Jordi Grau-Moya, Li Kevin Wenliang, Matthew Aitchison, Laurent Orseau, et al. 2023 · 2023
Cited alongside, same era.
Llama 3: A family of large language models
Meta AI. 2024 · 2024
Closest in time.
Reducing transformer key-value cache size with cross-layer attention
William Brandon, Mayank Mishra, Aniruddha Nrusimha, Rameswar Panda, and Jonathan Ragan Kelly. 2024 · 2024
Closest in time.
Pyramidkv: Dynamic kv cache compression based on pyramidal information funneling
Zefan Cai, Yichi Zhang, Bofei Gao, Yuliang Liu, Tianyu Liu, Keming Lu, Wayne Xiong, Yue Dong, Baobao Chang, Junjie Hu, et al. 2024 · 2024
Closest in time.
Kvquant: Towards 10 million context length llm inference with kv cache quantization
Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami. 2024 · 2024
Closest in time.
Ruler: What’s the real context size of your long-context language models?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Suyu Ge, Yunan Zhang, Liyuan Liu, Minjia Zhang, Jiawei Han, and Jianfeng Gao. 2023 · 2023
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023 · 2023
Cited alongside, same era.
Compressing context to enhance inference efficiency of large language models
Yucheng Li, Bo Dong, Chenghua Lin, and Frank Guerin. 2023 · 2023
Cited alongside, same era.
Deja vu: Contextual sparsity for efficient llms at inference time
Zichang Liu, Jue Wang, Tri Dao, Tianyi Zhou, Binhang Yuan, Zhao Song, Anshumali Shrivastava, Ce Zhang, Yuandong Tian, Christopher Re, et al. 2023 · 2023
Cited alongside, same era.
Attention sorting combats recency bias in long context language models
Alexander Peysakhovich and Adam Lerer. 2023 · 2023
Cited alongside, same era.
Efficiently scaling transformer inference
Reiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Jonathan Heek, Kefan Xiao, Shivani Agrawal, and Jeff Dean. 2023 · 2023
Cited alongside, same era.
The truth is in there: Improving reasoning in language models with layer-selective rank reduction
Pratyusha Sharma, Jordan T Ash, and Dipendra Misra. 2023 · 2023
Cited alongside, same era.
Flexgen: High-throughput generative inference of large language models with a single gpu
Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Ré, Ion Stoica, and Ce Zhang. 2023 · 2023
Cited alongside, same era.
Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg. 2024 · 2024
Closest in time.
Compression represents intelligence linearly
Yuzhen Huang, Jinghan Zhang, Zifei Shan, and Junxian He. 2024 · 2024
Closest in time.
Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention
Huiqiang Jiang, Yucheng Li, Chengruidong Zhang, Qianhui Wu, Xufang Luo, Surin Ahn, Zhenhua Han, Amir H Abdi, Dongsheng Li, Chin-Yew Lin, et al. 2024 · 2024
Closest in time.
Snapkv: Llm knows what you are looking for before generation
Yuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh, Acyr Locatelli, Hanchen Ye, Tianle Cai, Patrick Lewis, and Deming Chen. 2024 · 2024
Closest in time.
Massive activations in large language models
Mingjie Sun, Xinlei Chen, J Zico Kolter, and Zhuang Liu. 2024 · 2024
Closest in time.
Scaling laws with vocabulary: Larger models deserve larger vocabularies
Chaofan Tao, Qian Liu, Longxu Dou, Niklas Muennighoff, Zhongwei Wan, Ping Luo, Min Lin, and Ngai Wong. 2024 · 2024
Closest in time.
Attention is all you need but you don’t need all of it for inference of large language models
Georgy Tyukin, Gbetondji JS Dovonon, Jean Kaddour, and Pasquale Minervini. 2024 · 2024
Closest in time.
D2o: Dynamic discriminative operations for efficient generative inference of large language models
Zhongwei Wan, Xinjian Wu, Yu Zhang, Yi Xin, Chaofan Tao, Zhihong Zhu, Xin Wang, Siqi Luo, Jing Xiong, and Mi Zhang. 2024 · 2024
Closest in time.
Retrieval head mechanistically explains long-context factuality
Wenhao Wu, Yizhong Wang, Guangxuan Xiao, Hao Peng, and Yao Fu. 2024 · 2024
Closest in time.
Duoattention: Efficient long-context llm inference with retrieval and streaming heads
Guangxuan Xiao, Jiaming Tang, Jingwei Zuo, Junxian Guo, Shang Yang, Haotian Tang, Yao Fu, and Song Han. 2024 · 2024
Closest in time.
Cacheblend: Fast large language model serving with cached knowledge fusion
Jiayi Yao, Hanchen Li, Yuhan Liu, Siddhant Ray, Yihua Cheng, Qizheng Zhang, Kuntai Du, Shan Lu, and Junchen Jiang. 2024 · 2024
Closest in time.
Effectively compress kv heads for llm
Hao Yu, Zelan Yang, Shen Li, Yong Li, and Jianxin Wu. 2024 · 2024
Closest in time.
Prepacking: A simple method for fast prefilling and increased throughput in large language models
Siyan Zhao, Daniel Israel, Guy Van den Broeck, and Aditya Grover. 2024 · 2024
Closest in time.
Dape v2: Process attention score as feature map for length extrapolation
Chuanyang Zheng, Yihang Gao, Han Shi, Jing Xiong, Jiankai Sun, Jingyao Li, Minbin Huang, Xiaozhe Ren, Michael Ng, Xin Jiang, et al. 2024 · 2024
Closest in time.
Parallelcomp: Parallel long-context compressor for length extrapolation
Jing Xiong, Jianghan Shen, Chuanyang Zheng, Zhongwei Wan, Chenyang Zhao, Chiwun Yang, Fanghua Ye, Hongxia Yang, Lingpeng Kong, and Ngai Wong. 2025 · 2025
Closest in time.
Chuanyang Zheng, Yihang Gao, Guoxuan Chen, Han Shi, Jing Xiong, Xiaozhe Ren, Chao Huang, Xin Jiang, Zhenguo Li, and Yu Li. 2025 · 2025
Closest in time.