Fetching the paper…
Reading the bibliography…
The growing demand for efficient Large Language Model (LLM) inference requires a holistic optimization on algorithms, systems, and hardware.
Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Earlier work this paper cites.
Multi-news: a large-scale multi-document summarization dataset and abstractive hierarchical model, 2019
A. R. Fabbri, I. Li, T. She, S. Li, and D. R. Radev · 2019
Earlier work this paper cites.
Decoupled weight decay regularization, 2019
I. Loshchilov and F. Hutter · 2019
Earlier work this paper cites.
Program synthesis with large language models, 2021
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, and C. Sutton · 2021
Earlier work this paper cites.
Evaluating large language models trained on code, 2021
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, and W. Zaremba · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Earlier work this paper cites.
Efficient attentions for long document summarization, 2021
L. Huang, S. Cao, N. Parulian, H. Ji, and L. Wang · 2021
Earlier work this paper cites.
Perceiver: General perception with iterative attention, 2021
A. Jaegle, F. Gimeno, A. Brock, A. Zisserman, O. Vinyals, and J. Carreira · 2021
Earlier work this paper cites.
Language models are multilingual chain-of-thought reasoners, 2022
F. Shi, M. Suzgun, M. Freitag, X. Wang, S. Srivats, S. Vosoughi, H. W. Chung, Y. Tay, S. Ruder, D. Zhou, D. Das, and J. Wei · 2022
Earlier work this paper cites.
Accelerating large language model decoding with speculative sampling, 2023
C. Chen, S. Borgeaud, G. Irving, J.-B. Lespiau, L. Sifre, and J. Jumper · 2023
Earlier work this paper cites.
Fast inference from transformers via speculative decoding, 2023
Y. Leviathan, M. Kalman, and Y. Matias · 2023
Earlier work this paper cites.
Holistic evaluation of language models, 2023
P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y. Zhang, D. Narayanan, Y. Wu, A. Kumar, B. Newman, B. Yuan, B. Yan, C. Zhang, C. Cosgrove, C. D. Manning, C. Ré, D. Acosta-Navas, D. A. Hudson, E. Zelikman, E. Durmus, F. Ladhak, F. Rong, H. Ren, H. Yao, J. Wang, K. Santhanam, L. Orr, L. Zheng, M. Yuksekgonul, M. Suzgun, N. Kim, N. Guha, N. Chatterji, O. Khattab, P. Henderson, Q. Huang, R. Chi, S. M. Xie, S. Santurkar, S. Ganguli, T. Hashimoto, T. Icard, T. Zhang, V. Chaudhary, W. Wang, X. Li, Y. Mai, Y. Zhang, and Y. Koreeda · 2023
Earlier work this paper cites.
Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation
J. Liu, C. S. Xia, Y. Wang, and L. Zhang · 2023
Earlier work this paper cites.
Towards efficient generative large language model serving: A survey from algorithms to systems, 2023
X. Miao, G. Oliaro, Z. Zhang, X. Cheng, H. Jin, T. Chen, and Z. Jia · 2023
Earlier work this paper cites.
Megabyte: Predicting million-byte sequences with multiscale transformers, 2023
L. Yu, D. Simig, C. Flaherty, A. Aghajanyan, L. Zettlemoyer, and M. Lewis · 2023
Cited alongside, same era.
H 2 o: Heavy-hitter oracle for efficient generative inference of large language models, 2023
Z. Zhang, Y. Sheng, T. Zhou, T. Chen, L. Zheng, R. Cai, Z. Song, Y. Tian, C. Ré, C. Barrett, Z. Wang, and B. Chen · 2023
Cited alongside, same era.
Instruction-following evaluation for large language models, 2023
J. Zhou, T. Lu, S. Mishra, S. Brahma, S. Basu, Y. Luan, D. Zhou, and L. Hou · 2023
Cited alongside, same era.
Longbench: A bilingual, multitask benchmark for long context understanding, 2024
Y. Bai, X. Lv, J. Zhang, H. Lyu, J. Tang, Z. Huang, Z. Du, X. Liu, A. Zeng, L. Hou, Y. Dong, J. Tang, and J. Li · 2024
Cited alongside, same era.
Medusa: Simple llm inference acceleration framework with multiple decoding heads, 2024
T. Cai, Y. Li, Z. Geng, H. Peng, J. D. Lee, D. Chen, and T. Dao · 2024
Cited alongside, same era.
Byte latent transformer: Patches scale better than tokens, 2024
A. Pagnoni, R. Pasunuru, P. Rodriguez, J. Nguyen, B. Muller, M. Li, C. Zhou, L. Yu, J. Weston, L. Zettlemoyer, G. Ghosh, M. Lewis, A. Holtzman, and S. Iyer · 2024
Later among the works it cites.
Discovering the gems in early layers: Accelerating long-context llms with 1000x input token reduction, 2024
Z. Shi, Y. Ming, X.-P. Nguyen, Y. Liang, and S. Joty · 2024
Later among the works it cites.
Quest: Query-aware sparsity for efficient long-context llm inference, 2024
J. Tang, Y. Zhao, K. Zhu, G. Xiao, B. Kasikci, and S. Han · 2024
Later among the works it cites.
Large concept models: Language modeling in a sentence representation space, 2024
L. team, L. Barrault, P.-A. Duquenne, M. Elbayad, A. Kozhevnikov, B. Alastruey, P. Andrews, M. Coria, G. Couairon, M. R. Costa-jussà, D. Dale, H. Elsahar, K. Heffernan, J. M. Janeiro, T. Tran, C. Ropers, E. Sánchez, R. S. Roman, A. Mourachko, S. Saleem, and H. Schwenk · 2024
Later among the works it cites.
Unlocking efficiency in large language model inference: A comprehensive survey of speculative decoding, 2024
H. Xia, Z. Yang, Q. Dong, P. Wang, Y. Li, T. Ge, T. Liu, W. Li, and Z. Sui · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Magicpig: Lsh sampling for efficient llm generation, 2024
Z. Chen, R. Sadhukhan, Z. Ye, Y. Zhou, J. Zhang, N. Nolte, Y. Tian, M. Douze, L. Bottou, Z. Jia, and B. Chen · 2024
Cited alongside, same era.
How far are we from agi: Are llms all we need?, 2024
T. Feng, C. Jin, J. Liu, K. Zhu, H. Tu, Z. Cheng, G. Lin, and J. You · 2024
Cited alongside, same era.
The language model evaluation harness, 07 2024
L. Gao, J. Tow, B. Abbasi, S. Biderman, S. Black, A. DiPofi, C. Foster, L. Golding, J. Hsu, A. Le Noac’h, H. Li, K. McDonell, N. Muennighoff, C. Ociepa, J. Phang, L. Reynolds, H. Schoelkopf, A. Skowron, L. Sutawika, E. Tang, A. Thite, B. Wang, K. Wang, and A. Zou · 2024
Cited alongside, same era.
Block transformer: Global-to-local language modeling for fast inference
N. Ho, S. Bae, T. Kim, H. Jo, Y. Kim, T. Schuster, A. Fisch, J. Thorne, and S.-Y. Yun · 2024
Cited alongside, same era.
Kvquant: Towards 10 million context length llm inference with kv cache quantization, 2024
C. Hooper, S. Kim, H. Mohammadzadeh, M. W. Mahoney, Y. S. Shao, K. Keutzer, and A. Gholami · 2024
Cited alongside, same era.
Snapkv: Llm knows what you are looking for before generation, 2024
Y. Li, Y. Huang, B. Yang, B. Venkitesh, A. Locatelli, H. Ye, T. Cai, P. Lewis, and D. Chen · 2024
Cited alongside, same era.
Chunk-distilled language modeling, 2024
Y. Li, K. Livescu, and J. Zhou · 2024
Cited alongside, same era.
Later among the works it cites.
Efficient streaming language models with attention sinks, 2024
G. Xiao, Y. Tian, B. Chen, S. Han, and M. Lewis · 2024
Later among the works it cites.
Dyspec: Faster speculative decoding with dynamic token tree structure, 2024
Y. Xiong, R. Zhang, Y. Li, T. Wu, and L. Zou · 2024
Later among the works it cites.
Lookahead: An inference acceleration framework for large language model with lossless generation accuracy, 2024
Y. Zhao, Z. Xie, C. Liang, C. Zhuang, and J. Gu · 2024
Later among the works it cites.
The claude 3 model family: Opus, sonnet, haiku
Anthropic · 2025
Closest in time.
Eagle: Speculative sampling requires rethinking feature uncertainty, 2025
Y. Li, F. Wei, C. Zhang, and H. Zhang · 2025
Closest in time.
Speculative prefill: Turbocharging ttft with lightweight and training-free token importance estimation, 2025
J. Liu, B. Chen, and C. Zhang · 2025
Closest in time.
Qwen2.5 technical report, 2025
Qwen, :, A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, and Z. Qiu · 2025
Closest in time.
Shadowkv: Kv cache in shadows for high-throughput long-context llm inference, 2025
H. Sun, L.-W. Chang, W. Bao, S. Zheng, N. Zheng, X. Liu, H. Dong, Y. Chi, and B. Chen · 2025
Closest in time.
Deft: Decoding with flash tree-attention for efficient tree-structured llm inference, 2025
J. Yao, K. Chen, K. Zhang, J. You, B. Yuan, Z. Wang, and T. Lin · 2025
Closest in time.