Fetching the paper…
Reading the bibliography…
Large language models have been widely deployed in various applications, encompassing both interactive online tasks and batched offline tasks.
Managing flash crowds on the internet
2003
Earlier work this paper cites.
Network traffic characteristics of data centers in the wild
2010
Earlier work this paper cites.
Characterizing, modeling, and generating workload spikes for stateful services
2010
Earlier work this paper cites.
Ara: Adaptive resource allocation for cloud computing environments under bursty workloads
2011
Earlier work this paper cites.
Measuring cloud workload burstiness
2014
Earlier work this paper cites.
Evaluating the impact of fine-scale burstiness on cloud elasticity
2015
Earlier work this paper cites.
Large-scale cluster management at google with borg
2015
Earlier work this paper cites.
Self-adaptive resource allocation for energy-aware virtual machine placement in dynamic computing cloud
2018
Earlier work this paper cites.
How many words do we read per minute? a review and meta-analysis of reading rate
2019
Earlier work this paper cites.
Shenango: Achieving high CPU efficiency for latency-sensitive datacenter workloads
2019
Earlier work this paper cites.
Fast and exact analysis for lru caches
2019
Earlier work this paper cites.
Burstable instances for clouds: Performance modeling, equilibrium analysis, and revenue maximization
2020
Earlier work this paper cites.
What makes good in-context examples for gpt-
2021
Earlier work this paper cites.
Next-qa:next phase of question-answering to explaining temporal actions, 2021
2021
Earlier work this paper cites.
Can foundation models wrangle your data?, 2022
2022
Earlier work this paper cites.
Sarathi: Efficient llm inference by piggybacking decodes with chunked prefills, 2023
2023
Earlier work this paper cites.
Sharegpt
2023
Earlier work this paper cites.
Large language models are zero-shot rankers for recommender systems
2023
Earlier work this paper cites.
Evaluating open-domain question answering in the era of large language models
2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
2023
Earlier work this paper cites.
Holistic evaluation of language models, 2023
2023
Earlier work this paper cites.
Splitwise: Efficient generative llm inference using phase splitting
2023
Earlier work this paper cites.
Efficiently scaling transformer inference
2023
Earlier work this paper cites.
Code llama: Open foundation models for code
2023
Earlier work this paper cites.
Flexgen: High-throughput generative inference of large language models with a single gpu, 2023
2023
Cited alongside, same era.
Autogpt: Build, deploy, and run ai agents
2023
Cited alongside, same era.
Llm-planner: Few-shot grounded planning for embodied agents with large language models, 2023
2023
Cited alongside, same era.
Llama: Open and efficient foundation language models, 2023
2023
Cited alongside, same era.
Gödel: Unified large-scale resource management and scheduling at bytedance
2023
Cited alongside, same era.
Phi-4 technical report, 2024
2024
Cited alongside, same era.
Spotserve: Serving generative large language models on preemptible instances
2024
Later among the works it cites.
Learning to reason with llms
2024
Later among the works it cites.
Openai batch api
2024
Later among the works it cites.
Conserve: Harvesting gpus for low-latency and high-throughput large language model serving, 2024
2024
Later among the works it cites.
Mooncake: A kvcache-centric disaggregated architecture for llm serving, 2024
2024
Later among the works it cites.
Qwq: Reflect deeply on the boundaries of the unknown
2024
Later among the works it cites.
Preble: Efficient distributed prompt scheduling for llm serving, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2024
Cited alongside, same era.
Introducing the message batches api
2024
Cited alongside, same era.
Moe-lightning: High-throughput moe inference on memory-constrained gpus, 2024
2024
Cited alongside, same era.
Efficient and economic large language model inference with attention offloading, 2024
2024
Cited alongside, same era.
Chatbot arena: An open platform for evaluating llms by human preference, 2024
2024
Cited alongside, same era.
Cost-Efficient large language model serving for multi-turn conversations with CachedAttention
2024
Cited alongside, same era.
2024
Later among the works it cites.
Large language models for data annotation and synthesis: A survey, 2024
2024
Later among the works it cites.
[core] adding priority scheduling
2024
Later among the works it cites.
vllm automatic prefix caching
2024
Later among the works it cites.
Yi-lightning technical report, 2024
2024
Later among the works it cites.
Revisiting slo and goodput metrics in llm serving, 2024
2024
Later among the works it cites.
Loongserve: Efficiently serving long-context large language models with elastic sequence parallelism
2024
Later among the works it cites.
Inference scaling laws: An empirical analysis of compute-optimal inference for problem-solving with language models, 2024
2024
Later among the works it cites.
Pie: Pooling cpu memory for llm inference
2024
Later among the works it cites.
Benchmarking large language models for news summarization
2024
Later among the works it cites.
Blendserve: Optimizing offline inference for auto-regressive large models with resource-aware batching, 2024
2024
Later among the works it cites.
Recommender systems in the era of large language models (llms), 2024
2024
Later among the works it cites.
Sglang: Efficient execution of structured language model programs, 2024
2024
Later among the works it cites.
Batchllm: Optimizing large batched llm inference with global prefix sharing and throughput-oriented token batching, 2024
2024
Later among the works it cites.
Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving
2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
2025
Closest in time.
Qwen2.5 technical report, 2025
2025
Closest in time.