Fetching the paper…
Reading the bibliography…
Striking an optimal balance between minimal drafting latency and high speculation accuracy to enhance the inference speed of Large Language Models remains a significant challenge in speculative decoding.
Fine-Tuning by Curriculum Learning for Non-Autoregressive Neural Machine Translation
Guo, J.; Tan, X.; Xu, L.; Qin, T.; Chen, E.; and Liu, T.-Y. 2019b · 1911
Earlier work this paper cites.
Coding and information theory , volume 134
Roman, S. 1992 · 1992
Earlier work this paper cites.
Long short-term memory
Hochreiter, S.; and Schmidhuber, J. 1997 · 1997
Earlier work this paper cites.
Glancing Transformer for Non-Autoregressive Neural Machine Translation
Qian, L.; Zhou, H.; Bao, Y.; Wang, M.; Qiu, L.; Zhang, W.; Yu, Y.; and Li, L. 2021 · 2003
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Graves, A.; Fernández, S.; Gomez, F.; and Schmidhuber, J. 2006 · 2006
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Hinton, G.; Vinyals, O.; and Dean, J. 2015 · 2015
Earlier work this paper cites.
WaveNet: A Generative Model for Raw Audio
van den Oord, A.; Dieleman, S.; Zen, H.; Simonyan, K.; Vinyals, O.; Graves, A.; Kalchbrenner, N.; Senior, A.; and Kavukcuoglu, K. 2016 · 2016
Earlier work this paper cites.
Blockwise Parallel Decoding for Deep Autoregressive Models
Stern, M.; Shazeer, N.; and Uszkoreit, J. 2018 · 2018
Earlier work this paper cites.
Parallel WaveNet: Fast High-Fidelity Speech Synthesis
van den Oord, A.; Li, Y.; Babuschkin, I.; Simonyan, K.; Vinyals, O.; Kavukcuoglu, K.; van den Driessche, G.; Lockhart, E.; Cobo, L.; Stimberg, F.; Casagrande, N.; Grewe, D.; Noury, S.; Dieleman, S.; Elsen, E.; Kalchbrenner, N.; Zen, H.; Graves, A.; King, H.; Walters, T.; Belov, D.; and Hassabis, D. 2018 · 2018
Earlier work this paper cites.
Semi-Autoregressive Neural Machine Translation
Wang, C.; Zhang, J.; and Chen, H. 2018 · 2018
Earlier work this paper cites.
Fast Decoding in Sequence Models using Discrete Latent Variables
Łukasz Kaiser; Roy, A.; Vaswani, A.; Parmar, N.; Bengio, S.; Uszkoreit, J.; and Shazeer, N. 2018 · 2018
Earlier work this paper cites.
Mask-Predict: Parallel Decoding of Conditional Masked Language Models
Ghazvininejad, M.; Levy, O.; Liu, Y.; and Zettlemoyer, L. 2019 · 2019
Earlier work this paper cites.
Fast Structured Decoding for Sequence Models
Sun, Z.; Li, Z.; Wang, H.; He, D.; Lin, Z.; and Deng, Z. 2019 · 2019
Earlier work this paper cites.
Non-Autoregressive Machine Translation with Auxiliary Regularization
Wang, Y.; Tian, F.; He, D.; Qin, T.; Zhai, C.; and Liu, T.-Y. 2019 · 2019
Earlier work this paper cites.
Fine-Tuning by Curriculum Learning for Non-Autoregressive Neural Machine Translation
Guo, J.; Tan, X.; Xu, L.; Qin, T.; Chen, E.; and Liu, T.-Y. 2020 · 2020
Earlier work this paper cites.
Learning to Recover from Multi-Modality Errors for Non-Autoregressive Neural Machine Translation
Ran, Q.; Lin, Y.; Li, P.; and Zhou, J. 2020 · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M.; Tworek, J.; Jun, H.; Yuan, Q.; Pinto, H. P. D. O.; Kaplan, J.; Edwards, H.; Burda, Y.; Joseph, N.; Brockman, G.; et al. 2021 · 2021
Cited alongside, same era.
Training verifiers to solve math word problems
Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; et al. 2021 · 2021
Cited alongside, same era.
Understanding and Improving Lexical Choice in Non-Autoregressive Translation
Ding, L.; Wang, L.; Liu, X.; Wong, D. F.; Tao, D.; and Tu, Z. 2021 · 2021
Cited alongside, same era.
Learning to Rewrite for Non-Autoregressive Neural Machine Translation
Geng, X.; Feng, X.; and Qin, B. 2021 · 2021
Cited alongside, same era.
Self-Distillation Mixup Training for Non-autoregressive Neural Machine Translation
Accelerating LLM Inference with Staged Speculative Decoding
Spector, B.; and Re, C. 2023 · 2023
Later among the works it cites.
Speculative Decoding: Lossless Speedup of Autoregressive Translation
Xia, H.; Ge, T.; Chen, S.-Q.; Wei, F.; and Sui, Z. 2023 · 2023
Later among the works it cites.
A Survey on Non-Autoregressive Generation for Neural Machine Translation and Beyond
Xiao, Y.; Wu, L.; Guo, J.; Li, J.; Zhang, M.; Qin, T.; and yan Liu, T. 2023 · 2023
Later among the works it cites.
Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
Cai, T.; Li, Y.; Geng, Z.; Peng, H.; Lee, J. D.; Chen, D.; and Dao, T. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Guo, J.; Wang, M.; Wei, D.; Shang, H.; Wang, Y.; Li, Z.; Yu, Z.; Wu, Z.; Chen, Y.; Su, C.; Zhang, M.; Lei, L.; shimin tao; and Yang, H. 2021 · 2021
Cited alongside, same era.
How Does Distilled Data Complexity Impact the Quality and Confidence of Non-Autoregressive Machine Translation?
Xu, W.; Ma, S.; Zhang, D.; and Carpuat, M. 2021 · 2021
Cited alongside, same era.
latent-GLAT: Glancing at Latent Variables for Parallel Text Generation
Bao, Y.; Zhou, H.; Huang, S.; Wang, D.; Qian, L.; Dai, X.; Chen, J.; and Li, L. 2022 · 2022
Cited alongside, same era.
GLM: General Language Model Pretraining with Autoregressive Blank Infilling
Du, Z.; Qian, Y.; Liu, X.; Ding, M.; Qiu, J.; Yang, Z.; and Tang, J. 2022 · 2022
Cited alongside, same era.
In-context Learning and Induction Heads
Olsson, C.; Elhage, N.; Nanda, N.; Joseph, N.; DasSarma, N.; Henighan, T.; Mann, B.; Askell, A.; Bai, Y.; Chen, A.; Conerly, T.; Drain, D.; Ganguli, D.; Hatfield-Dodds, Z.; Hernandez, D.; Johnston, S.; Jones, A.; Kernion, J.; Lovitt, L.; Ndousse, K.; Amodei, D.; Brown, T.; Clark, J.; Kaplan, J.; McCandlish, S.; and Olah, C. 2022 · 2022
Cited alongside, same era.
Accelerating Large Language Model Decoding with Speculative Sampling
Chen, C.; Borgeaud, S.; Irving, G.; Lespiau, J.-B.; Sifre, L.; and Jumper, J. 2023 · 2023
Cited alongside, same era.
Speculative Decoding with Big Little Decoder
Kim, S.; Mangalam, K.; Moon, S.; Malik, J.; Mahoney, M. W.; Gholami, A.; and Keutzer, K. 2023 · 2023
Cited alongside, same era.
Fast Inference from Transformers via Speculative Decoding
Leviathan, Y.; Kalman, M.; and Matias, Y. 2023 · 2023
Cited alongside, same era.
Chen, Z.; Yang, X.; Lin, J.; Sun, C.; Chang, K. C.-C.; and Huang, J. 2024 · 2024
Closest in time.
Better & Faster Large Language Models via Multi-token Prediction
Gloeckle, F.; Idrissi, B. Y.; Rozière, B.; Lopez-Paz, D.; and Synnaeve, G. 2024 · 2024
Closest in time.
EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
Li, Y.; Wei, F.; Zhang, C.; and Zhang, H. 2024 · 2024
Closest in time.
SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and Verification
Miao, X.; Oliaro, G.; Zhang, Z.; Cheng, X.; Wang, Z.; Zhang, Z.; Wong, R. Y. Y.; Zhu, A.; Yang, L.; Shi, X.; Shi, C.; Chen, Z.; Arfeen, D.; Abhyankar, R.; and Jia, Z. 2024 · 2024
Closest in time.
Efficient Large Language Models: A Survey
Wan, Z.; Wang, X.; Liu, C.; Alam, S.; Zheng, Y.; Liu, J.; Qu, Z.; Yan, S.; Zhu, Y.; Zhang, Q.; Chowdhury, M.; and Zhang, M. 2024 · 2024
Closest in time.
Accelerating Production LLMs with Combined Token/Embedding Speculators
Wertheimer, D.; Rosenkranz, J.; Parnell, T.; Suneja, S.; Ranganathan, P.; Ganti, R.; and Srivatsa, M. 2024 · 2024
Closest in time.
Xia, H.; Yang, Z.; Dong, Q.; Wang, P.; Li, Y.; Ge, T.; Liu, T.; Li, W.; and Sui, Z. 2024 · 2024
Closest in time.
Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding
Zhang, J.; Wang, J.; Li, H.; Shou, L.; Chen, K.; Chen, G.; and Mehrotra, S. 2024 · 2024
Closest in time.
Zhao, Y.; Xie, Z.; Liang, C.; Zhuang, C.; and Gu, J. 2024 · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.; et al. 2024 · 2024
Closest in time.
A Survey on Model Compression for Large Language Models
Zhu, X.; Li, J.; Liu, Y.; Ma, C.; and Wang, W. 2024 · 2024
Closest in time.