Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) present a critical trade-off between inference quality and computational cost: larger models offer superior capabilities but incur significant latency, while smaller models are faster but less powerful.
The exponentially weighted moving average
Hunter, J. S · 1986
Earlier work this paper cites.
Routing design in operational networks: A look from the inside
Maltz, D. A., Xie, G., Zhan, J., Zhang, H., Hjálmtỳsson, G., and Greenberg, A · 2004
Earlier work this paper cites.
Clipper: A { \{ Low-Latency } \} online prediction serving system
Crankshaw, D., Wang, X., Zhou, G., Franklin, M. J., Gonzalez, J. E., and Stoica, I · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A · 2019
Earlier work this paper cites.
Nvidia a100 gpu: Performance & innovation for gpu computing
Choquette, J., and Gandhi, W · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. D. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
Unified scaling laws for routed language models
Clark, A., de Las Casas, D., Guy, A., Mensch, A., Paganini, M., Hoffmann, J., Damoc, B., Hechtman, B., Cai, T., Borgeaud, S., et al · 2022
Earlier work this paper cites.
Language models are multilingual chain-of-thought reasoners
Shi, F., Suzgun, M., Freitag, M., Wang, X., Srivats, S., Vosoughi, S., Chung, H. W., Tay, Y., Ruder, S., Zhou, D., et al · 2022
Earlier work this paper cites.
Orca: A distributed serving system for { \{ Transformer-Based } \} generative models
Yu, G.-I., Jeong, J. S., Kim, G.-W., Kim, S., and Chun, B.-G · 2022
Earlier work this paper cites.
Scaling laws for generative mixed-modal language models
Aghajanyan, A., Yu, L., Conneau, A., Hsu, W.-N., Hambardzumyan, K., Zhang, S., Roller, S., Goyal, N., Levy, O., and Zettlemoyer, L · 2023
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., and Stoica, I · 2023
Earlier work this paper cites.
Efficient inference with model cascades
Lebovitz, L., Cavigelli, L., Magno, M., and Muller, L. K · 2023
Earlier work this paper cites.
Fast inference from transformers via speculative decoding
Leviathan, Y., Kalman, M., and Matias, Y · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Cited alongside, same era.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2023
Cited alongside, same era.
Vidur: A large-scale simulation framework for llm inference
Agrawal, A., Kedia, N., Mohan, J., Panwar, A., Kwatra, N., Gulavani, B. S., Ramjee, R., and Tumanov, A · 2024
Cited alongside, same era.
Medusa: Simple llm inference acceleration framework with multiple decoding heads
Cai, T., Li, Y., Geng, Z., Peng, H., Lee, J. D., Chen, D., and Dao, T · 2024
Cited alongside, same era.
Routellm: Learning to route llms from preference data
Ong, I., Almahairi, A., Wu, V., Chiang, W.-L., Wu, T., Gonzalez, J. E., Kadous, M. W., and Stoica, I · 2024
Later among the works it cites.
Powerinfer: Fast large language model serving with a consumer-grade gpu
Song, Y., Mi, Z., Xie, H., and Chen, H · 2024
Later among the works it cites.
Llumnix: Dynamic scheduling for large language model serving
Sun, B., Huang, Z., Zhao, H., Xiao, W., Zhang, X., Li, Y., and Lin, W · 2024
Later among the works it cites.
Loongserve: Efficiently serving long-context large language models with elastic sequence parallelism
Wu, B., Liu, S., Zhong, Y., Sun, P., Liu, X., and Jin, X · 2024
Later among the works it cites.
A theoretical perspective for speculative decoding algorithm
Yin, M., Chen, M., Huang, K., and Wang, M · 2024
Later among the works it cites.
Fast and live model auto scaling with o(1) host caching, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cascade speculative drafting for even faster llm inference
Chen, Z., Yang, X., Lin, J., Sun, C., Chang, K., and Huang, J · 2024
Cited alongside, same era.
Increasing transformer token length with a maximum entropy principle method
Cukier, R · 2024
Cited alongside, same era.
Graphrouter: A graph-based router for llm selections
Feng, T., Shen, Y., and You, J · 2024
Cited alongside, same era.
Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Vaughan, A., et al · 2024
Cited alongside, same era.
Characterization of large language model development in the datacenter
Hu, Q., Ye, Z., Wang, Z., Wang, G., Zhang, M., Chen, Q., Sun, P., Lin, D., Wang, X., Luo, Y., et al · 2024
Cited alongside, same era.
A taxonomy and survey on grid-based routing protocols designed for wireless sensor networks
Jain, S., and Verma, R. K · 2024
Cited alongside, same era.
EAGLE: Speculative sampling requires rethinking feature uncertainty
Li, Y., Wei, F., Zhang, C., and Zhang, H · 2024
Cited alongside, same era.
Zhang, D., Wang, H., Liu, Y., Wei, X., Shan, Y., Chen, R., and Chen, H · 2024
Later among the works it cites.
Sglang: Efficient execution of structured language model programs
Zheng, L., Yin, L., Xie, Z., Sun, C. L., Huang, J., Yu, C. H., Cao, S., Kozyrakis, C., Stoica, I., Gonzalez, J. E., et al · 2024
Later among the works it cites.
{ \{ DistServe } \} : Disaggregating prefill and decoding for goodput-optimized large language model serving
Zhong, Y., Liu, S., Chen, J., Hu, J., Zhu, Y., Liu, X., Jin, X., and Zhang, H · 2024
Later among the works it cites.
A survey on efficient inference for large language models
Zhou, Z., Ning, X., Hong, K., Fu, T., Xu, J., Li, S., Lou, Y., Wang, L., Yuan, Z., Li, X., et al · 2024
Later among the works it cites.
Specserve: Efficient and slo-aware large language model serving with adaptive speculative decoding
Huang, K., Wu, H., Shi, Z., Zou, H., Yu, M., and Shi, Q · 2025
Closest in time.
Static batching of irregular workloads on gpus: Framework and application to efficient moe model inference, 2025
Li, Y., Li, Y., Zhang, J., Chen, B., Chen, X., Duan, L., Jin, Y., Li, Z., Liu, X., Wang, H., Wang, W., Wang, Y., Yang, J., Zhang, P., Zheng, L., and Yu, W · 2025
Closest in time.
EAGLE-3: Scaling up inference acceleration of large language models via training-time test, 2025
Li, Y., Wei, F., Zhang, C., and Zhang, H · 2025
Closest in time.
Adaserve: Slo-customized llm serving with fine-grained speculative decoding
Li, Z., Chen, Z., Delacourt, R., Oliaro, G., Wang, Z., Chen, Q., Lin, S., Yang, A., Zhang, Z., Chen, Z., et al · 2025
Closest in time.
Pearl: Parallel speculative decoding with adaptive draft length
Liu, T., Li, Y., Lv, Q., Liu, K., Zhu, J., Hu, W., and Sun, X · 2025
Closest in time.
Leveraging uncertainty estimation for efficient llm routing
Zhang, T., Mehradfar, A., Dimitriadis, D., and Avestimehr, S · 2025
Closest in time.
Fastcache: Optimizing multimodal llm serving through lightweight kv-cache compression framework
Zhu, J., Wu, H., Wang, H., Li, Y., Hou, B., Li, R., and Zhai, J · 2025
Closest in time.