Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have been driving a new wave of interactive AI applications across numerous domains.
Regression models with ordinal variables
Christopher Winship and Robert D Mare · 1984
Earlier work this paper cites.
LsPS: A job size-based scheduler for efficient task assignments in Hadoop
Yi Yao, Jianzhe Tai, Bo Sheng, and Ningfang Mi · 2015
Earlier work this paper cites.
Clipper: A low-latency online prediction serving system
Daniel Crankshaw, Xin Wang, Guilio Zhou, Michael J Franklin, Joseph E Gonzalez, and Ion Stoica · 2017
Earlier work this paper cites.
TensorFlow-Serving: Flexible, high-performance ML serving
Christopher Olston, Fangwei Li, Jeremiah Harmsen, Jordan Soyke, Kiril Gorovoy, Li Lao, Noah Fiedel, Sukriti Ramesh, and Vinu Rajashekhar · 2017
Earlier work this paper cites.
ZygOS: Achieving low tail latency for microsecond-scale networked tasks
George Prekas, Marios Kogias, and Edouard Bugnion · 2017
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional Transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
PipeSwitch: Fast pipelined context switching for deep learning applications
Zhihao Bai, Zhen Zhang, Yibo Zhu, and Xin Jin · 2020
Earlier work this paper cites.
InferLine: Latency-aware provisioning and scaling for prediction serving pipelines
Daniel Crankshaw, Gur-Eyal Sela, Xiangxi Mo, Corey Zumar, Ion Stoica, Joseph Gonzalez, and Alexey Tumanov · 2020
Earlier work this paper cites.
Serving DNNs like Clockwork: Performance predictability from the bottom up
Arpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao, Antoine Kaufmann, Ymir Vigfusson, and Jonathan Mace · 2020
Earlier work this paper cites.
A hybrid scheduling platform: a runtime prediction reliability aware scheduling platform to improve HPC scheduling performance
Mina Naghshnejad and Mukesh Singhal · 2020
Earlier work this paper cites.
Faster and cheaper serverless computing on harvested resources
Yanqi Zhang, Íñigo Goiri, Gohar Irfan Chaudhry, Rodrigo Fonseca, Sameh Elnikety, Christina Delimitrou, and Ricardo Bianchini · 2021
Cited alongside, same era.
PREP: Predicting job runtime with job running path on supercomputers
Longfang Zhou, Xiaorong Zhang, Wenxiang Yang, Yongguo Han, Fang Wang, Yadong Wu, and Jie Yu · 2021
Cited alongside, same era.
A case for task sampling based learning for cluster job scheduling
Akshay Jajoo, Y. Charlie Hu, Xiaojun Lin, and Nan Deng · 2022
Cited alongside, same era.
Fluid limits for shortest job first with aging
Yonatan Shadmi · 2022
Cited alongside, same era.
Orca: A distributed serving system for Transformer-based generative models
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun · 2022
Cited alongside, same era.
GPTCache: An open-source semantic cache for LLM applications enabling faster answers and cost savings
AlpaServe: Statistical multiplexing with model parallelism for deep learning serving
Zhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu, Ying Sheng, Xin Jin, Yanping Huang, Zhifeng Chen, Hao Zhang, Joseph E Gonzalez, et al · 2023
Later among the works it cites.
Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Rae Ying Yee Wong, Zhuoming Chen, Daiyaan Arfeen, Reyna Abhyankar, and Zhihao Jia · 2023
Later among the works it cites.
Fast distributed inference serving for large language models
Bingyang Wu, Yinmin Zhong, Zili Zhang, Gang Huang, Xuanzhe Liu, and Xin Jin · 2023
Later among the works it cites.
LMSYS-Chat-1M: A large-scale real-world LLM conversation dataset
Lianmin Zheng, Wei-Lin Chiang, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fu Bang and Feng Di · 2023
Cited alongside, same era.
Medusa: Simple framework for accelerating LLM generation with multiple decoding heads, 2023
Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, and Tri Dao · 2023
Cited alongside, same era.
Break the sequential dependency of LLM inference using lookahead decoding, 2023
Yichao Fu, Peter Bailis, Ion Stoica, and Hao Zhang · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with PagedAttention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica · 2023
Cited alongside, same era.
Fast inference from transformers via speculative decoding
Yaniv Leviathan, Matan Kalman, and Yossi Matias · 2023
Cited alongside, same era.
https://developer.nvidia.com/triton-inference-server
Triton inference server
Cited in the paper.
Zangwei Zheng, Xiaozhe Ren, Fuzhao Xue, Yang Luo, Xin Jiang, and Yang You · 2023
Later among the works it cites.
GitHub Copilot
GitHub · 2024
Closest in time.
An important next step on our AI journey
Google · 2024
Closest in time.
Inference without interference: Disaggregate LLM inference for mixed downstream workloads, 2024
Cunchen Hu, Heyang Huang, Liangliang Xu, Xusheng Chen, Jiang Xu, Shuang Chen, Hao Feng, Chenxi Wang, Sa Wang, Yungang Bao, Ninghui Sun, and Yizhou Shan · 2024
Closest in time.
Introducing ChatGPT
OpenAI · 2024
Closest in time.
https://github.com/vllm-project/vllm , 2024
vLLM: A high-throughput and memory-efficient inference and serving engine for LLMs · 2024
Closest in time.