Fetching the paper…
Reading the bibliography…
Large Language Model (LLM) workloads have distinct prefill and decode phases with different compute and memory requirements which should ideally be accounted for when scheduling input queries across different LLM instances in a cluster.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
Parallel data, tools and interfaces in OPUS
Tiedemann, J · 2012
Earlier work this paper cites.
User-priority guided min-min scheduling algorithm for load balancing in cloud computing
Chen, H., Wang, F., Helian, N., and Akanmu, G · 2013
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
Results of the WNUT2017 shared task on novel and emerging entity recognition
Derczynski, L., Nichols, E., van Erp, M., and Limsopatham, N · 2017
Earlier work this paper cites.
ELI5: long form question answering
Fan, A., Jernite, Y., Perez, E., Grangier, D., Weston, J., and Auli, M · 2019
Earlier work this paper cites.
Towards a human-like open-domain chatbot
Adiwardana, D., Luong, M.-T., So, D. R., Hall, J., Fiedel, N., Thoppilan, R., Yang, Z., Kulshreshtha, A., Nemade, G., Lu, Y., et al · 2020
Earlier work this paper cites.
Recipes for building an open-domain chatbot
Roller, S., Dinan, E., Goyal, N., Ju, D., Williamson, M., Liu, Y., Xu, J., Ott, M., Shuster, K., Smith, E. M., et al · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Earlier work this paper cites.
Heuristic-guided reinforcement learning
Cheng, C.-A., Kolobov, A., and Swaminathan, A · 2021
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D., Ermon, S., Rudra, A., and Ré, C · 2022
Earlier work this paper cites.
On efficient approximate queries over machine learning models
Ding, D., Amer-Yahia, S., and Lakshmanan, L. V · 2022
Earlier work this paper cites.
Efficient edge inference by selective query
Kag, A., Fedorov, I., Gangrade, A., Whatmough, P., and Saligrama, V · 2022
Earlier work this paper cites.
Orca: A distributed serving system for Transformer-Based generative models
Yu, G.-I., Jeong, J. S., Kim, G.-W., Kim, S., and Chun, B.-G · 2022
Earlier work this paper cites.
Sarathi: Efficient llm inference by piggybacking decodes with chunked prefills
Agrawal, A., Panwar, A., Mohan, J., Kwatra, N., Gulavani, B. S., and Ramjee, R · 2023
Earlier work this paper cites.
$s^3$: Increasing GPU utilization during generative inference for higher throughput
Jin, Y., Wu, C.-F., Brooks, D., and Wei, G.-Y · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., and Stoica, I · 2023
Cited alongside, same era.
Gpt-4 technical report, 2023
OpenAI · 2023
Cited alongside, same era.
Splitwise: Efficient generative llm inference using phase splitting, 2023
Patel, P., Choukse, E., Zhang, C., Íñigo Goiri, Shah, A., Maleki, S., and Bianchini, R · 2023
Cited alongside, same era.
Accelerating llm inference with staged speculative decoding
Spector, B. and Re, C · 2023
Cited alongside, same era.
Rlq: Workload allocation with reinforcement learning in distributed queues
Staffolani, A., Darvariu, V.-A., Bellavista, P., and Musolesi, M · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models, 2023
Blockllm: Multi-tenant finer-grained serving for large language models
Li, J., Xu, L., Xu, H., and Akella, A · 2024
Closest in time.
Infinite-llm: Efficient llm service for long context with distattention and distributed kvcache
Lin, B., Peng, T., Zhang, C., Sun, M., Li, L., Zhao, H., Xiao, W., Xu, Q., Qiu, X., Li, S., et al · 2024
Closest in time.
Andes: Defining and enhancing quality-of-experience in llm-based text streaming services
Liu, J., Wu, Z., Chung, J.-W., Lai, F., Lee, M., and Chowdhury, M · 2024
Closest in time.
Model selection for latency-critical inference serving
Mendoza, D., Romero, F., and Trippel, C · 2024
Closest in time.
Aladdin: Joint placement and scaling for slo-aware llm serving
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T · 2023
Cited alongside, same era.
Fast distributed inference serving for large language models
Wu, B., Zhong, Y., Zhang, Z., Huang, G., Liu, X., and Jin, X · 2023
Cited alongside, same era.
Taming throughput-latency tradeoff in llm inference with sarathi-serve, 2024
Agrawal, A., Kedia, N., Panwar, A., Mohan, J., Kwatra, N., Gulavani, B. S., Tumanov, A., and Ramjee, R · 2024
Cited alongside, same era.
Hybrid llm: Cost-efficient and quality-aware query routing
Ding, D., Mallick, A., Wang, C., Sim, R., Mukherjee, S., Ruhle, V., Lakshmanan, L. V., and Awadallah, A. H · 2024
Cited alongside, same era.
Efficient llm scheduling by learning to rank
Fu, Y., Zhu, S., Su, R., Qiao, A., Stoica, I., and Zhang, H · 2024
Cited alongside, same era.
Prompt cache: Modular attention reuse for low-latency inference
Gim, I., Chen, G., Lee, S.-s., Sarda, N., Khandelwal, A., and Zhong, L · 2024
Cited alongside, same era.
Inference without interference: Disaggregate llm inference for mixed downstream workloads
Hu, C., Huang, H., Xu, L., Chen, X., Xu, J., Chen, S., Feng, H., Wang, C., Wang, S., Bao, Y., et al · 2024
Cited alongside, same era.
Nie, C., Fonseca, R., and Liu, Z · 2024
Closest in time.
Routellm: Learning to route llms with preference data
Ong, I., Almahairi, A., Wu, V., Chiang, W.-L., Wu, T., Gonzalez, J. E., Kadous, M. W., and Stoica, I · 2024
Closest in time.
One queue is all you need: Resolving head-of-line blocking in large language model serving, 2024
Patke, A., Reddy, D., Jha, S., Qiu, H., Pinto, C., Cui, S., Narayanaswami, C., Kalbarczyk, Z., and Iyer, R · 2024
Closest in time.
vattention: Dynamic memory management for serving llms without pagedattention
Prabhu, R., Nayak, A., Mohan, J., Ramjee, R., and Panwar, A · 2024
Closest in time.
Efficient interactive llm serving with proxy model-based sequence length prediction
Qiu, H., Mao, W., Patke, A., Cui, S., Jha, S., Wang, C., Franke, H., Kalbarczyk, Z. T., Başar, T., and Iyer, R. K · 2024
Closest in time.
Skippredict: When to invest in predictions for scheduling
Shahout, R. and Mitzenmacher, M · 2024
Closest in time.
Don’t stop me now: Embedding based scheduling for llms
Shahout, R., Malach, E., Liu, C., Jiang, W., Yu, M., and Mitzenmacher, M · 2024
Closest in time.
Llumnix: Dynamic scheduling for large language model serving
Sun, B., Huang, Z., Zhao, H., Xiao, W., Zhang, X., Li, Y., and Lin, W · 2024
Closest in time.
H2o: Heavy-hitter oracle for efficient generative inference of large language models
Zhang, Z., Sheng, Y., Zhou, T., Chen, T., Zheng, L., Cai, R., Song, Z., Tian, Y., Ré, C., Barrett, C., et al · 2024
Closest in time.
Response length perception and sequence scheduling: An llm-empowered llm inference pipeline
Zheng, Z., Ren, X., Xue, F., Luo, Y., Jiang, X., and You, Y · 2024
Closest in time.
Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving
Zhong, Y., Liu, S., Chen, J., Hu, J., Zhu, Y., Liu, X., Jin, X., and Zhang, H · 2024
Closest in time.