Fetching the paper…
Reading the bibliography…
This technical report describes the design and training of novel speculative decoding draft models, for accelerating the inference speeds of large language models in a production environment.
Pytorch compile to speed up inference on llama 2
Antoni Viros i Martin, Brian Vaughan, Davis Wertheimer, Joshua Rosenkranz, Mudhakar Srivatsa, Nelson Mimura Gonzalez, Raghu Ganti, Supriyo Chakraborty, Zhuoran Liu, Geeta Chauhan, and Hamid Shojanazeri · 2023
Earlier work this paper cites.
Fast inference from transformers via speculative decoding
Yaniv Leviathan, Matan Kalman, and Yossi Matias · 2023
Earlier work this paper cites.
Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Zhengxin Zhang, Rae Ying Yee Wong, Alan Zhu, Lijie Yang, Xiaoxiang Shi, Chunan Shi, Zhuoming Chen, Daiyaan Arfeen, Reyna Abhyankar, and Zhihao Jia · 2023
Earlier work this paper cites.
Hydra: Sequentially-dependent draft heads for medusa decoding
Zachary Ankner, Rishab Parthasarathy, Aniruddha Nrusimha, Chistopher Rinard, Jonathan Ragan-Kelley, and William Brandon · 2024
Earlier work this paper cites.
Maximizing training throughput using pytorch fsdp
Team PyTorch at IBM and Team PyTorch at Meta · 2024
Cited alongside, same era.
Medusa: Simple llm inference acceleration framework with multiple decoding heads
Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D. Lee, Deming Chen, and Tri Dao · 2024
Cited alongside, same era.
Break the sequential dependency of llm inference using lookahead decoding
Yichao Fu, Peter Bailis, Ion Stoica, and Hao Zhang · 2024
Cited alongside, same era.
Eagle: Speculative sampling requires rethinking feature uncertainty
Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang · 2024
Closest in time.
Multi-candidate speculative decoding
Sen Yang, Shujian Huang, Xinyu Dai, and Jiajun Chen · 2024
Closest in time.
Recurrent drafter for fast speculative decoding in large language models
Aonan Zhang, Chong Wang, Yi Wang, Xuanyu Zhang, and Yunfei Cheng · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…