Fetching the paper…
Reading the bibliography…
We introduce CopySpec, a simple yet effective technique to tackle the inefficiencies LLMs face when generating responses that closely resemble previous outputs or responses that can be verbatim extracted from context.
Pointer networks
Oriol Vinyals, Meire Fortunato, and Navdeep Jaitly. 2015 · 2015
Earlier work this paper cites.
Incorporating copying mechanism in sequence-to-sequence learning
Jiatao Gu, Zhaopeng Lu, Hang Li, and Victor OK Li. 2016 · 2016
Earlier work this paper cites.
Get to the point: Summarization with pointer-generator networks
Abigail See, Peter J. Liu, and Christopher D. Manning. 2017 · 2017
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, and 39 others. 2021 · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021 · 2021
Earlier work this paper cites.
Parallel decoding with speculative sampling for large language models
Jia Chen and Hao Xu. 2023 · 2023
Earlier work this paper cites.
Speed: Speculative pipelined execution for efficient decoding in large language models
Zhang He and Xin Wang. 2023 · 2023
Earlier work this paper cites.
Rest: Retrieval-based speculative decoding
Zhenyu He and 1 others. 2023 · 2023
Earlier work this paper cites.
Fast inference from transformers via speculative decoding
Yossi Leviathan and 1 others. 2023 · 2023
Earlier work this paper cites.
Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation
Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. 2023 · 2023
Cited alongside, same era.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark. 2023 · 2023
Cited alongside, same era.
Xupeng Miao and 1 others. 2023 · 2023
Cited alongside, same era.
Prompt lookup decoding
Apoorv Saxena. 2023 · 2023
Cited alongside, same era.
Predictive pipelined decoding: A compute-latency trade-off for exact llm decoding
Break the sequential dependency of llm inference using lookahead decoding
Yichao Fu, Peter Bailis, Ion Stoica, and Hao Zhang. 2024 · 2024
Later among the works it cites.
Aaron Grattafiori and 1 others. 2024 · 2024
Later among the works it cites.
Sam decoding: Speculative decoding via suffix automaton
Yuxuan Hu, Ke Wang, Xiaokang Zhang, Fanjin Zhang, Cuiping Li, Hong Chen, and Jing Zhang. 2024 · 2024
Later among the works it cites.
Adaptive draft-verification for efficient large language model decoding
Xukun Liu, Yifan Zhang, Peiyi Wang, Tao Ge, Tianyu Liu, Yongqi Li, and Zhifang Sui. 2024 · 2024
Later among the works it cites.
Turning trash into treasure: Accelerating inference of large language models with token recycling
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Seongjun Yang and 1 others. 2023 · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 · 2023
Cited alongside, same era.
Distillspec: Improving speculative decoding via knowledge distillation
Yongchao Zhou and 1 others. 2023 · 2023
Cited alongside, same era.
Medusa: Multidraft speculative decoding for accelerated inference
Yu Cai and 1 others. 2024 · 2024
Cited alongside, same era.
Hardware-aware parallel prompt decoding for memory-efficient acceleration of llm inference
Hao Mark Chen, Wayne Luk, Ka Fai Cedric Yiu, Rui Li, Konstantin Mishchenko, Stylianos I. Venieris, and Hongxiang Fan. 2024 · 2024
Cited alongside, same era.
Xianzhen Luo, Yixuan Wang, Qingfu Zhu, Zhiming Zhang, Xuanyu Zhang, Qing Yang, Dongliang Xu, and Wanxiang Che. 2024 · 2024
Later among the works it cites.
Pld+: Accelerating llm inference by leveraging language model artifacts
Shwetha Somasundaram, Anirudh Phukan, and Apoorv Saxena. 2024 · 2024
Later among the works it cites.
Spechub: Provable acceleration to multi-draft speculative decoding
Ryan Sun, Tianyi Zhou, Xun Chen, and Lichao Sun. 2024 · 2024
Later among the works it cites.
Eagle: Speculative sampling requires rethinking feature uncertainty
Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang. 2025 · 2025
Closest in time.
Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, and 25 others. 2025 · 2025
Closest in time.