Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have shown outstanding performance across numerous real-world tasks.
The state of sparsity in deep neural networks
Gale, T.; Elsen, E.; and Hooker, S. 2019 · 1902
Earlier work this paper cites.
Mask-predict: Parallel decoding of conditional masked language models
Ghazvininejad, M.; Levy, O.; Liu, Y.; and Zettlemoyer, L. 2019 · 1904
Earlier work this paper cites.
Tree structured analysis on GPU power study
Chen, J.; Li, B.; Zhang, Y.; Peng, L.; and Peir, J.-k. 2011 · 2011
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G.; Vinyals, O.; and Dean, J. 2015 · 2015
Earlier work this paper cites.
Characterizing power and performance of gpu memory access
Allen, T.; and Ge, R. 2016 · 2016
Earlier work this paper cites.
Non-autoregressive neural machine translation
Gu, J.; Bradbury, J.; Xiong, C.; Li, V. O.; and Socher, R. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Deterministic non-autoregressive neural sequence modeling by iterative refinement
Lee, J.; Mansimov, E.; and Cho, K. 2018 · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for SQuAD
Rajpurkar, P.; Jia, R.; and Liang, P. 2018 · 2018
Earlier work this paper cites.
Yu, T.; Zhang, R.; Yang, K.; Yasunaga, M.; Wang, D.; Li, Z.; Ma, J.; Li, I.; Yao, Q.; Roman, S.; et al. 2018 · 2018
Earlier work this paper cites.
Fast structured decoding for sequence models
Sun, Z.; Li, Z.; Wang, H.; He, D.; Lin, Z.; and Deng, Z. 2019 · 2019
Earlier work this paper cites.
Non-autoregressive machine translation with auxiliary regularization
Wang, Y.; Tian, F.; He, D.; Qin, T.; Zhai, C.; and Liu, T.-Y. 2019 · 2019
Earlier work this paper cites.
Jointly masked sequence-to-sequence model for non-autoregressive neural machine translation
Guo, J.; Xu, L.; and Chen, E. 2020 · 2020
Earlier work this paper cites.
Movement pruning: Adaptive sparsity by fine-tuning
Sanh, V.; Wolf, T.; and Rush, A. 2020 · 2020
Cited alongside, same era.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Frantar, E.; Ashkboos, S.; Hoefler, T.; and Alistarh, D. 2022 · 2022
Cited alongside, same era.
Opt: Open pre-trained transformer language models
Zhang, S.; Roller, S.; Goyal, N.; Artetxe, M.; Chen, M.; Chen, S.; Dewan, C.; Diab, M.; Li, X.; Lin, X. V.; et al. 2022 · 2022
Cited alongside, same era.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Cited alongside, same era.
Spectr: Fast speculative decoding via optimal transport
Sun, Z.; Suresh, A. T.; Ro, J. H.; Beirami, A.; Jain, H.; and Yu, F. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Later among the works it cites.
Part-time Power Measurements: nvidia-smi’s Lack of Attention
Yang, Z.; Adamek, K.; and Armour, W. 2023 · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.; et al. 2023 · 2023
Later among the works it cites.
Llama 3 Model Card
AI@Meta. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Anil, R.; Dai, A. M.; Firat, O.; Johnson, M.; Lepikhin, D.; Passos, A.; Shakeri, S.; Taropa, E.; Bailey, P.; Chen, Z.; et al. 2023 · 2023
Cited alongside, same era.
Accelerating large language model decoding with speculative sampling
Chen, C.; Borgeaud, S.; Irving, G.; Lespiau, J.-B.; Sifre, L.; and Jumper, J. 2023 · 2023
Cited alongside, same era.
Rest: Retrieval-based speculative decoding
He, Z.; Zhong, Z.; Cai, T.; Lee, J. D.; and He, D. 2023 · 2023
Cited alongside, same era.
Speculative decoding with big little decoder
Kim, S.; Mangalam, K.; Moon, S.; Malik, J.; Mahoney, M. W.; Gholami, A.; and Keutzer, K. 2023 · 2023
Cited alongside, same era.
Fast inference from transformers via speculative decoding
Leviathan, Y.; Kalman, M.; and Matias, Y. 2023 · 2023
Cited alongside, same era.
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Lin, J.; Tang, J.; Tang, H.; Yang, S.; Dang, X.; and Han, S. 2023 · 2023
Cited alongside, same era.
Liu, X.; Hu, L.; Bailis, P.; Stoica, I.; Deng, Z.; Cheung, A.; and Zhang, H. 2023 · 2023
Cited alongside, same era.
Miao, X.; Oliaro, G.; Zhang, Z.; Cheng, X.; Wang, Z.; Wong, R. Y. Y.; Chen, Z.; Arfeen, D.; Abhyankar, R.; and Jia, Z. 2023 · 2023
Cited alongside, same era.
Andronov, M.; Andronova, N.; Wand, M.; Schmidhuber, J.; and Clevert, D.-A. 2024 · 2024
Closest in time.
Medusa: Simple llm inference acceleration framework with multiple decoding heads
Cai, T.; Li, Y.; Geng, Z.; Peng, H.; Lee, J. D.; Chen, D.; and Dao, T. 2024 · 2024
Closest in time.
Gear: An efficient kv cache compression recipefor near-lossless generative inference of llm
Kang, H.; Zhang, Q.; Kundu, S.; Jeong, G.; Liu, Z.; Krishna, T.; and Zhao, T. 2024 · 2024
Closest in time.
EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees
Li, Y.; Wei, F.; Zhang, C.; and Zhang, H. 2024 · 2024
Closest in time.
Dynamic-Width Speculative Beam Decoding for Efficient LLM Inference
Qin, Z.; He, Z.; Prakriya, N.; Cong, J.; and Sun, Y. 2024 · 2024
Closest in time.
A thorough examination of decoding methods in the era of llms
Shi, C.; Yang, H.; Cai, D.; Zhang, Z.; Wang, Y.; Yang, Y.; and Lam, W. 2024 · 2024
Closest in time.
Multi-candidate speculative decoding
Yang, S.; Huang, S.; Dai, X.; and Chen, J. 2024 · 2024
Closest in time.