Fetching the paper…
Reading the bibliography…
Decoding with autoregressive large language models (LLMs) traditionally occurs sequentially, generating one token after another.
Pytorch: An imperative style, high-performance deep learning library, 2019
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Köpf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 1912
Earlier work this paper cites.
How not to lie with statistics: the correct way to summarize benchmark results
Fleming, P. J. and Wallace, J. J · 1986
Earlier work this paper cites.
Scaling laws for neural language models, 2020
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2001
Earlier work this paper cites.
Alter: exploiting breakable dependences for parallelization
Udupa, A., Rajan, K., and Thies, W · 2011
Earlier work this paper cites.
Dancing with uncertainty
Misailovic, S., Sidiroglou, S., and Rinard, M. C · 2012
Earlier work this paper cites.
Blockwise parallel decoding for deep autoregressive models
Stern, M., Shazeer, N., and Uszkoreit, J · 2018
Earlier work this paper cites.
Xla : Compiling machine learning for peak performance, 2020
Sabne, A · 2020
Earlier work this paper cites.
Reducing activation recomputation in large transformer models, 2022
Korthikanti, V., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B · 2022
Earlier work this paper cites.
Efficiently scaling transformer inference, 2022
Pope, R., Douglas, S., Chowdhery, A., Devlin, J., Bradbury, J., Levskaya, A., Heek, J., Xiao, K., Agrawal, S., and Dean, J · 2022
Earlier work this paper cites.
Accelerating large language model decoding with speculative sampling
Chen, C., Borgeaud, S., Irving, G., Lespiau, J.-B., Sifre, L., and Jumper, J · 2023
Earlier work this paper cites.
Rest: Retrieval-based speculative decoding
He, Z., Zhong, Z., Cai, T., Lee, J. D., and He, D · 2023
Earlier work this paper cites.
Fast inference from transformers via speculative decoding
Leviathan, Y., Kalman, M., and Matias, Y · 2023
Earlier work this paper cites.
Alpacaeval: An automatic evaluator of instruction-following models
Li, X., Zhang, T., Dubois, Y., Taori, R., Gulrajani, I., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Earlier work this paper cites.
Taskmatrix.ai: Completing tasks by connecting foundation models with millions of apis, 2023
Liang, Y., Wu, C., Song, T., Wu, W., Xia, Y., Liu, Y., Ou, Y., Lu, S., Ji, L., Mao, S., Wang, Y., Shou, L., Gong, M., and Duan, N · 2023
Cited alongside, same era.
Chameleon: Plug-and-play compositional reasoning with large language models, 2023
Lu, P., Peng, B., Cheng, H., Galley, M., Chang, K.-W., Wu, Y. N., Zhu, S.-C., and Gao, J · 2023
Cited alongside, same era.
Skeleton-of-thought: Large language models can do parallel decoding
Ning, X., Lin, Z., Zhou, Z., Yang, H., and Wang, Y · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2023
Cited alongside, same era.
Accelerating transformer inference for translation via parallel decoding
Hydra: Sequentially-dependent draft heads for medusa decoding
Ankner, Z., Parthasarathy, R., Nrusimha, A., Rinard, C., Ragan-Kelley, J., and Brandon, W · 2024
Later among the works it cites.
A general theoretical paradigm to understand learning from human preferences
Azar, M. G., Guo, Z. D., Piot, B., Munos, R., Rowland, M., Valko, M., and Calandriello, D · 2024
Later among the works it cites.
Medusa: Simple llm inference acceleration framework with multiple decoding heads
Cai, T., Li, Y., Geng, Z., Peng, H., Lee, J. D., Chen, D., and Dao, T · 2024
Later among the works it cites.
Length-controlled alpacaeval: A simple way to debias automatic evaluators
Dubois, Y., Galambosi, B., Liang, P., and Hashimoto, T. B · 2024
Later among the works it cites.
Break the sequential dependency of llm inference using lookahead decoding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Santilli, A., Severino, S., Postolache, E., Maiorca, V., Mancusi, M., Marin, R., and Rodolà, E · 2023
Cited alongside, same era.
Toolformer: Language models can teach themselves to use tools, 2023
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., and Scialom, T · 2023
Cited alongside, same era.
Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face, 2023
Shen, Y., Song, K., Tan, X., Li, D., Lu, W., and Zhuang, Y · 2023
Cited alongside, same era.
Accelerating llm inference with staged speculative decoding
Spector, B. and Re, C · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models, 2023
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Cited alongside, same era.
React: Synergizing reasoning and acting in language models, 2023
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y · 2023
Cited alongside, same era.
Efficiently programming large language models using sglang
Zheng, L., Yin, L., Xie, Z., Huang, J., Sun, C., Yu, C. H., Cao, S., Kozyrakis, C., Stoica, I., Gonzalez, J. E., et al · 2023
Cited alongside, same era.
Gemini: A family of highly capable multimodal models, 2024
Anil, R., Borgeaud, S., and et al., J.-B. A · 2024
Cited alongside, same era.
Fu, Y., Bailis, P., Stoica, I., and Zhang, H · 2024
Later among the works it cites.
Bonbon alignment for large language models and the sweetness of best-of-n sampling, 2024
Gui, L., Gârbacea, C., and Veitch, V · 2024
Later among the works it cites.
Jaech, A., Kalai, A., and et al., A. L · 2024
Later among the works it cites.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., de las Casas, D., Hanna, E. B., Bressand, F., Lengyel, G., Bour, G., Lample, G., Lavaud, L. R., Saulnier, L., Lachaux, M.-A., Stock, P., Subramanian, S., Yang, S., Antoniak, S., Scao, T. L., Gervet, T., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2024
Later among the works it cites.
Apar: Llms can do auto-parallel auto-regressive decoding
Liu, M., Zeng, A., Wang, B., Zhang, P., Tang, J., and Dong, Y · 2024
Later among the works it cites.
Gemma: Open models based on gemini research and technology, 2024
Mesnard, T., Hardin, C., and et al., R. D · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
Daya Guo, Dejian Yang, H. Z. e. a · 2025
Closest in time.
gpt-fast: High-performance gpt decoding
Liang, Y., Feng, B., He, H., and et al · 2025
Closest in time.