Fetching the paper…
Reading the bibliography…
Sequential dependencies present a fundamental bottleneck in deploying large-scale autoregressive models, particularly for real-time applications.
2016
Earlier work this paper cites.
S. Teerapittayanon, B. McDanel, and H. Kung, “Branchynet: Fast inference via early exiting from deep neural networks,” in 2016 23rd international conference on pattern recognition (ICPR) . IEEE, 2016, pp. 2464–2469
2016
Earlier work this paper cites.
A. Radford, “Improving language understanding by generative pre-training,” 2018
2018
Earlier work this paper cites.
S. Elfwing, E. Uchibe, and K. Doya, “Sigmoid-weighted linear units for neural network function approximation in reinforcement learning,” Neural networks , vol. 107, pp. 3–11, 2018
2018
Earlier work this paper cites.
M. Stern, N. Shazeer, and J. Uszkoreit, “Blockwise parallel decoding for deep autoregressive models,” Advances in Neural Information Processing Systems , vol. 31, 2018
2018
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
A. Slivkins et al. , “Introduction to multi-armed bandits,” Foundations and Trends® in Machine Learning , vol. 12, no. 1-2, pp. 1–286, 2019
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” Advances in neural information processing systems , vol. 33, pp. 6840–6851, 2020
2020
Earlier work this paper cites.
H. Liu, Z. Dai, D. So, and Q. V. Le, “Pay attention to mlps,” Advances in neural information processing systems , vol. 34, pp. 9204–9215, 2021
2021
Earlier work this paper cites.
M. Ding, Z. Yang, W. Hong, W. Zheng, C. Zhou, D. Yin, J. Lin, X. Zou, Z. Shao, H. Yang et al. , “Cogview: Mastering text-to-image generation via transformers,” Advances in neural information processing systems , vol. 34, pp. 19 822–19 835, 2021
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
M. U. Hadi, R. Qureshi, A. Shah, M. Irfan, A. Zafar, M. Shaikh, N. Akhtar, J. Wu, and S. Mirjalili, “A survey on large language models: Applications, challenges, limitations, and practical usage,” TechRxiv , 2023
2023
Earlier work this paper cites.
H. Huang, S. Wu, X. Liang, B. Wang, Y. Shi, P. Wu, M. Yang, and T. Zhao, “Towards making the most of llm for translation quality estimation,” in CCF International Conference on Natural Language Processing and Chinese Computing . Springer, 2023, pp. 375–386
2023
Earlier work this paper cites.
L. Yang, Z. Zhang, Y. Song, S. Hong, R. Xu, Y. Zhao, W. Zhang, B. Cui, and M.-H. Yang, “Diffusion models: A comprehensive survey of methods and applications,” ACM Computing Surveys , vol. 56, no. 4, pp. 1–39, 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
J. Z. Wu, Y. Ge, X. Wang, S. W. Lei, Y. Gu, Y. Shi, W. Hsu, Y. Shan, X. Qie, and M. Z. Shou, “Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 7623–7633
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
H. Xia, T. Ge, P. Wang, S.-Q. Chen, F. Wei, and Z. Sui, “Speculative decoding: Exploiting speculative execution for accelerating seq2seq generation,” in Findings of the Association for Computational Linguistics: EMNLP 2023 , 2023, pp. 3909–3925
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Y. Leviathan, M. Kalman, and Y. Matias, “Fast inference from transformers via speculative decoding,” in International Conference on Machine Learning . PMLR, 2023, pp. 19 274–19 286
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2024
Earlier work this paper cites.
“Open-sora report v1.1,” https://github.com/hpcaitech/Open-Sora/blob/main/docs/report_02.md , 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
D. Xu, W. Yin, H. Zhang, X. Jin, Y. Zhang, S. Wei, M. Xu, and X. Liu, “Edgellm: Fast on-device llm inference with speculative decoding,” IEEE Transactions on Mobile Computing , 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
W. Zhao, Y. Huang, X. Han, W. Xu, C. Xiao, X. Zhang, Y. Fang, K. Zhang, Z. Liu, and M. Sun, “Ouroboros: Generating longer drafts phrase by phrase for faster speculative decoding,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , Y. Al-Onaizan, M. Bansal, and Y.-N. Chen, Eds. Miami, Florida, USA: Association for Computational Linguistics, Nov. 2024, pp. 13 378–13 393. [Online]. Available: https://aclanthology.org/2024.emnlp-main.742/
2024
Later among the works it cites.
2024
Later among the works it cites.
H. Yi, F. Lin, H. Li, N. Peiyang, X. Yu, and R. Xiao, “Generation meets verification: Accelerating large language model inference with smart parallel auto-correct decoding,” in Findings of the Association for Computational Linguistics: ACL 2024 , L.-W. Ku, A. Martins, and V. Srikumar, Eds. Bangkok, Thailand: Association for Computational Linguistics, Aug. 2024, pp. 5285–5299. [Online]. Available: https://aclanthology.org/2024.findings-acl.313/
2024
Later among the works it cites.
2024
Later among the works it cites.
C. Li, Z. Zhou, S. Zheng, J. Zhang, Y. Liang, and G. Sun, “Specpim: Accelerating speculative inference on pim-enabled system via architecture-dataflow co-exploration,” in Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3 , 2024, pp. 950–965
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
B. Butler, S. Yu, A. Mazaheri, and A. Jannesari, “Pipeinfer: Accelerating llm inference using asynchronous pipelined speculation,” in SC24: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 2024, pp. 1–19
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
J. Liu, B. Park, and X. Shen, “A drop-in solution for on-the-fly adaptation of speculative decoding in large language models,” 2025. [Online]. Available: https://openreview.net/forum?id=xOtOfdbBqK
2025
Closest in time.
G. Bachmann, S. Anagnostidis, A. Pumarola, M. Georgopoulos, A. Sanakoyeu, Y. Du, E. Schönfeld, A. Thabet, and J. K. Kohler, “Judge decoding: Faster speculative sampling requires going beyond model alignment,” in The Thirteenth International Conference on Learning Representations , 2025. [Online]. Available: https://openreview.net/forum?id=mtSSFiqW6y
2025
Closest in time.
G. Liu, A. Ramachandran, T. Gangwani, Y. Fu, and A. Sethy, “Knowledge distillation with training wheels,” 2025. [Online]. Available: https://www.amazon.science/publications/knowledge-distillation-with-training-wheels
2025
Closest in time.
Z. Sun, U. Mendlovic, Y. Leviathan, A. Aharoni, A. Beirami, J. H. Ro, and A. T. Suresh, “Block verification accelerates speculative decoding,” in The Thirteenth International Conference on Learning Representations , 2025. [Online]. Available: https://openreview.net/forum?id=frsg32u0rO
2025
Closest in time.
Z. Hu, T. Zheng, V. Viswanathan, Z. Chen, R. A. Rossi, Y. Wu, D. Manocha, and H. Huang, “Towards optimal multi-draft speculative decoding,” in The Thirteenth International Conference on Learning Representations , 2025. [Online]. Available: https://openreview.net/forum?id=9KxnxWOBA5
2025
Closest in time.