Fetching the paper…
Reading the bibliography…
Recent advancements in Large Language Models (LLMs) have led to increasingly diverse requests, accompanied with varying resource (compute and memory) demands to serve them.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, D., Chen, M., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., et al · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 2019
Earlier work this paper cites.
Orca: A distributed serving system for { \{ Transformer-Based } \} generative models
Yu, G.-I., Jeong, J. S., Kim, G.-W., Kim, S., and Chun, B.-G · 2022
Earlier work this paper cites.
{ \{ PetS } \} : A unified framework for { \{ Parameter-Efficient } \} transformers serving
Zhou, Z., Wei, X., Zhang, J., and Sun, G · 2022
Earlier work this paper cites.
Petals: Collaborative inference and fine-tuning of large models
Borzunov, A., Baranchuk, D., Dettmers, T., Riabinin, M., Belkada, Y., Chumachenko, A., Samygin, P., and Raffel, C · 2023
Earlier work this paper cites.
Large language models in education: A focus on the complementary relationship between human teachers and chatgpt
Jeon, J. and Lee, S · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Earlier work this paper cites.
{ \{ AlpaServe } \} : Statistical multiplexing with model parallelism for deep learning serving
Li, Z., Zheng, L., Zhong, Y., Liu, V., Sheng, Y., Jin, X., Huang, Y., Chen, Z., Zhang, H., Gonzalez, J. E., et al · 2023
Earlier work this paper cites.
Deja vu: Contextual sparsity for efficient llms at inference time
Liu, Z., Wang, J., Dao, T., Zhou, T., Yuan, B., Song, Z., Shrivastava, A., Zhang, C., Tian, Y., Re, C., et al · 2023
Earlier work this paper cites.
A study of generative large language model for medical research and healthcare
Peng, C., Yang, X., Chen, A., Smith, K. E., PourNejatian, N., Costa, A. B., Martin, C., Flores, M. G., Zhang, Y., Magoc, T., et al · 2023
Earlier work this paper cites.
Fast distributed inference serving for large language models
Wu, B., Zhong, Y., Zhang, Z., Huang, G., Liu, X., and Jin, X · 2023
Earlier work this paper cites.
The claude 3 model family: Opus, sonnet, haiku, 2024
Anthropic · 2024
Earlier work this paper cites.
Azure public dataset, 2024
Azure · 2024
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Cited alongside, same era.
The world’s most widely adopted ai developer tool, 2024
GitHub · 2024
Cited alongside, same era.
M \ \backslash ’elange: Cost efficient large language model serving by exploiting gpu heterogeneity
Griggs, T., Liu, X., Yu, J., Kim, D., Chiang, W.-L., Cheung, A., and Stoica, I · 2024
Cited alongside, same era.
Inference without interference: Disaggregate llm inference for mixed downstream workloads
Hu, C., Huang, H., Xu, L., Chen, X., Xu, J., Chen, S., Feng, H., Wang, C., Wang, S., Bao, Y., et al · 2024
Cited alongside, same era.
Conserve: Harvesting gpus for low-latency and high-throughput large language model serving
Qiao, Y., Anzai, S., Yu, S., Ma, H., Wang, Y., Kim, M., and Xu, H · 2024
Later among the works it cites.
Mooncake: Kimi’s kvcache-centric architecture for llm serving
Qin, R., Li, Z., He, W., Zhang, M., Wu, Y., Zheng, W., and Xu, X · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M., Savinov, N., Teplyashin, D., Lepikhin, D., Lillicrap, T., Alayrac, J.-b., Soricut, R., Lazaridou, A., Firat, O., Schrittwieser, J., et al · 2024
Later among the works it cites.
Dynamollm: Designing llm inference clusters for performance and energy efficiency
Stojkovic, J., Zhang, C., Goiri, Í., Torrellas, J., and Choukse, E · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al · 2024
Cited alongside, same era.
Helix: Distributed serving of large language models via max-flow on heterogeneous gpus
Mei, Y., Zhuang, Y., Miao, X., Yang, J., Jia, Z., and Vinayak, R · 2024
Cited alongside, same era.
Spotserve: Serving generative large language models on preemptible instances
Miao, X., Shi, C., Duan, J., Xi, X., Lin, D., Cui, B., and Jia, Z · 2024
Cited alongside, same era.
Exegpt: Constraint-aware resource scheduling for llm inference
Oh, H., Kim, K., Kim, J., Kim, S., Lee, J., Chang, D.-s., and Seo, J · 2024
Cited alongside, same era.
Openai gpt-4o, 2024
OpenAI · 2024
Cited alongside, same era.
Splitwise: Efficient generative llm inference using phase splitting
Patel, P., Choukse, E., Zhang, C., Shah, A., Goiri, Í., Maleki, S., and Bianchini, R · 2024
Cited alongside, same era.
Queue management for slo-oriented large language model serving
Patke, A., Reddy, D., Jha, S., Qiu, H., Pinto, C., Narayanaswami, C., Kalbarczyk, Z., and Iyer, R · 2024
Cited alongside, same era.
Llumnix: Dynamic scheduling for large language model serving
Sun, B., Huang, Z., Zhao, H., Xiao, W., Zhang, X., Li, Y., and Lin, W · 2024
Later among the works it cites.
Flashflex: Accommodating large language model training over heterogeneous environment
Yan, R., Jiang, Y., Tao, W., Nie, X., Cui, B., and Yuan, B · 2024
Later among the works it cites.
Llm-pq: Serving llm on heterogeneous clusters with phase-aware partition and adaptive quantization
Zhao, J., Wan, B., Peng, Y., Lin, H., and Wu, C · 2024
Later among the works it cites.
{ \{ DistServe } \} : Disaggregating prefill and decoding for goodput-optimized large language model serving
Zhong, Y., Liu, S., Chen, J., Hu, J., Zhu, Y., Liu, X., Jin, X., and Zhang, H · 2024
Later among the works it cites.
Cuasmrl: Optimizing gpu sass schedules via deep reinforcement learning
He, G. and Yoneki, E · 2025
Closest in time.
Li, H., Fu, F., Ge, H., Lin, S., Wang, X., Niu, J., Miao, X., and Cui, B · 2025
Closest in time.
Hexgen-text2sql: Optimizing llm inference request scheduling for agentic text-to-sql workflow
Peng, Y., Jiang, Y., Wang, C., and Yuan, B · 2025
Closest in time.