Fetching the paper…
Reading the bibliography…
Serving systems for Large Language Models (LLMs) improve throughput by processing several requests concurrently.
Clipper: A Low-Latency online prediction serving system
2017
Earlier work this paper cites.
Ray: A distributed framework for emerging AI applications
2018
Earlier work this paper cites.
Learning scheduling algorithms for data processing clusters
2019
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
2020
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter, 2020
2020
Earlier work this paper cites.
Serverless in the wild: Characterizing and optimizing the serverless workload at a large cloud provider
2020
Earlier work this paper cites.
INFaaS: Automated model-less inference serving
2021
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models
2022
Earlier work this paper cites.
Optimizing mixture of experts using dynamic recompilations, 2022
2022
Earlier work this paper cites.
Orca: A distributed serving system for Transformer-Based generative models
2022
Earlier work this paper cites.
S3: Increasing GPU utilization during generative inference for higher throughput
2023
Earlier work this paper cites.
Extract-transform-load for video streams
2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
2023
Earlier work this paper cites.
S-lora: Serving thousands of concurrent lora adapters
2023
Earlier work this paper cites.
Fairness in serving large language models
2023
Cited alongside, same era.
Flexgen: high-throughput generative inference of large language models with a single gpu
2023
Cited alongside, same era.
Fast distributed inference serving for large language models, 2023
2023
Cited alongside, same era.
SHEPHERD: Serving DNNs in the wild
2023
Cited alongside, same era.
Response length perception and sequence scheduling: An llm-empowered llm inference pipeline
2023
Cited alongside, same era.
Etalon: Holistic performance evaluation framework for llm inference systems, 2024
2024
Cited alongside, same era.
Cascadeserve: Unlocking model cascades for inference serving, 2024
2024
Closest in time.
KServe Documentation
2024
Closest in time.
Kubernetes Documentation
2024
Closest in time.
MLFlow Documentation
2024
Closest in time.
Splitwise: Efficient generative llm inference using phase splitting
2024
Closest in time.
Don’t stop me now: Embedding based scheduling for llms, 2024
2024
Closest in time.
Skippredict: When to invest in predictions for scheduling, 2024
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Taming throughout-latency tradeoff in llm inference with sarathi-serve
2024
Cited alongside, same era.
Envoy Documentation
2024
Cited alongside, same era.
Envoy Load Balancer Documentation
2024
Cited alongside, same era.
Efficient llm scheduling by learning to rank, 2024
2024
Cited alongside, same era.
Istio Documentation
2024
Cited alongside, same era.
Istio Load Balancer Documentation
2024
Cited alongside, same era.
Llumnix: Dynamic scheduling for large language model serving
2024
Closest in time.
TensorRT-LLM Repository
2024
Closest in time.
Burstgpt: A real-world workload dataset to optimize llm serving systems, 2024
2024
Closest in time.
Stage: Query execution time prediction in amazon redshift
2024
Closest in time.
Blueprinting the cloud: Unifying and automatically optimizing cloud data infrastructures with brad
2024
Closest in time.
Sglang: Efficient execution of structured language model programs, 2024
2024
Closest in time.