Fetching the paper…
Reading the bibliography…
Large Language Model (LLM) collaborative decoding techniques improve output quality by combining the outputs of multiple models at each generation step, but they incur high computational costs.
Get to the point: Summarization with pointer-generator networks
See, A., Liu, P. J., and Manning, C. D · 2017
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Opt: Open pre-trained transformer language models, 2022
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., Mihaylov, T., Ott, M., Shleifer, S., Shuster, K., Simig, D., Koura, P. S., Sridhar, A., Wang, T., and Zettlemoyer, L · 2022
Earlier work this paper cites.
Accelerating large language model decoding with speculative sampling
Chen, C., Borgeaud, S., Irving, G., Lespiau, J.-B., Sifre, L., and Jumper, J · 2023
Earlier work this paper cites.
Fast inference from transformers via speculative decoding
Leviathan, Y., Kalman, M., and Matias, Y · 2023
Earlier work this paper cites.
Contrastive decoding: Open-ended text generation as optimization
Li, X. L., Holtzman, A., Fried, D., Liang, P., Eisner, J., Hashimoto, T. B., Zettlemoyer, L., and Lewis, M · 2023
Earlier work this paper cites.
Pass: Parallel speculative sampling
Monea, G., Joulin, A., and Grave, E · 2023
Earlier work this paper cites.
Contrastive decoding improves reasoning in large language models
O’Brien, S. and Lewis, M · 2023
Earlier work this paper cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Earlier work this paper cites.
Speculative decoding: Exploiting speculative execution for accelerating seq2seq generation
Xia, H., Ge, T., Wang, P., Chen, S.-Q., Wei, F., and Sui, Z · 2023
Earlier work this paper cites.
Judging llm-as-a-judge with mt-bench and chatbot arena, 2023
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I · 2023
Earlier work this paper cites.
Medusa: Simple llm inference acceleration framework with multiple decoding heads
Cai, T., Li, Y., Geng, Z., Peng, H., Lee, J. D., Chen, D., and Dao, T · 2024
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
Layerskip: Enabling early exit inference and self-speculative decoding
Elhoushi, M., Shrivastava, A., Liskovich, D., Hosmer, B., Wasti, B., Lai, L., Mahmoud, A., Acun, B., Agarwal, S., Roman, A., Aly, A., Chen, B., and Wu, C.-J · 2024
Cited alongside, same era.
Graph-structured speculative decoding
Gong, Z., Liu, J., Wang, Z., Wu, P., Wang, J., Cai, X., Zhao, D., and Yan, R · 2024
Cited alongside, same era.
Hybridmind: Meta selection of natural language and symbolic language for enhanced llm reasoning
Han, S., Liu, T., Li, C., Xiong, X., and Cohan, A · 2024
Cited alongside, same era.
Qwen2.5: A party of foundation models, September 2024
Team, Q · 2024
Later among the works it cites.
Mllm can see? dynamic correction decoding for hallucination mitigation
Wang, C., Chen, X., Zhang, N., Tian, B., Xu, H., Deng, S., and Chen, H · 2024
Later among the works it cites.
Determine-then-ensemble: Necessity of top-k union for large language model ensembling
Yao, Y., Wu, H., Liu, M., Luo, S., Han, X., Liu, J., Guo, Z., and Song, L · 2024
Later among the works it cites.
Yi, H., Lin, F., Li, H., Ning, P., Yu, X., and Xiao, R · 2024
Later among the works it cites.
Breaking the ceiling of the LLM community by treating token generation as a classification for ensembling
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Accelerated speculative sampling based on tree monte carlo
Hu, Z. and Huang, H · 2024
Cited alongside, same era.
Ensemble learning for heterogeneous large language models with deep parallel collaboration
Huang, Y., Feng, X., Li, B., Xiang, Y., Wang, H., Liu, T., and Qin, B · 2024
Cited alongside, same era.
Eagle: Speculative sampling requires rethinking feature uncertainty
Li, Y., Wei, F., Zhang, C., and Zhang, H · 2024
Cited alongside, same era.
Eagle-2: Faster inference of language models with dynamic draft trees
Li, Y., Wei, F., Zhang, C., and Zhang, H · 2024
Cited alongside, same era.
Decoding-time realignment of language models
Liu, T., Guo, S., Bianco, L., Calandriello, D., Berthet, Q., Llinares, F., Hoffmann, J., Dixon, L., Valko, M., and Blondel, M · 2024
Cited alongside, same era.
Lu, J., Pang, Z., Xiao, M., Zhu, Y., Xia, R., and Zhang, J · 2024
Cited alongside, same era.
Specinfer: Accelerating large language model serving with tree-based speculative inference and verification
Miao, X., Oliaro, G., Zhang, Z., Cheng, X., Wang, Z., Zhang, Z., Wong, R. Y. Y., Zhu, A., Yang, L., Shi, X., Shi, C., Chen, Z., Arfeen, D., Abhyankar, R., and Jia, Z · 2024
Cited alongside, same era.
Yu, Y.-C., Kuo, C. C., Ziqi, Y., Yucheng, C., and Li, Y.-S · 2024
Later among the works it cites.
Speculative contrastive decoding
Yuan, H., Lu, K., Huang, F., Yuan, Z., and Zhou, C · 2024
Later among the works it cites.
Draft&verify: Lossless large language model acceleration via self-speculative decoding
Zhang, J., Wang, J., Li, H., Shou, L., Chen, K., Chen, G., and Mehrotra, S · 2024
Later among the works it cites.
Distillspec: Improving speculative decoding via knowledge distillation
Zhou, Y., Lyu, K., Rawat, A. S., Menon, A., Rostamizadeh, A., Kumar, S., Kagy, J.-F., and Agarwal, R · 2024
Later among the works it cites.
Optimized multi-token joint decoding with auxiliary model for llm inference
Anonymous · 2025
Closest in time.
Judge decoding: Faster speculative sampling requires going beyond model alignment
Anonymous · 2025
Closest in time.
Harnessing multiple large language models: A survey on llm ensemble
Chen, Z., Li, J., Chen, P., Li, Z., Sun, K., Luo, Y., Mao, Q., Yang, D., Sun, H., and Yu, P. S · 2025
Closest in time.
Ateb: Evaluating and improving advanced nlp tasks for text embedding models
Han, S., Gomez, F. P., Vu, T., Li, Z., Cer, D., Zeng, H., Tar, C., Cohan, A., and Abrego, G. H · 2025
Closest in time.