Fetching the paper…
Reading the bibliography…
This paper presents ServerlessLLM, a distributed system designed to support low-latency serverless inference for Large Language Models (LLMs).
Zookeeper: Wait-free coordination for internet-scale systems
Patrick Hunt, Mahadev Konar, Flavio P. Junqueira, and Benjamin Reed · 2010
Earlier work this paper cites.
SAND: Towards High-Performance serverless computing
Istemi Ekin Akkus, Ruichuan Chen, Ivica Rimac, Manuel Stein, Klaus Satzke, Andre Beck, Paarijaat Aditya, and Volker Hilt · 2018
Earlier work this paper cites.
Pocket: Elastic ephemeral storage for serverless analytics
Ana Klimovic, Yawen Wang, Patrick Stuedi, Animesh Trivedi, Jonas Pfefferle, and Christos Kozyrakis · 2018
Earlier work this paper cites.
SOCK: Rapid task provisioning with Serverless-Optimized containers
Edward Oakes, Leon Yang, Dennis Zhou, Kevin Houck, Tyler Harter, Andrea Arpaci-Dusseau, and Remzi Arpaci-Dusseau · 2018
Earlier work this paper cites.
MArk: Exploiting cloud services for Cost-Effective, SLO-Aware machine learning inference serving
Chengliang Zhang, Minchen Yu, Wei Wang, and Feng Yan · 2019
Earlier work this paper cites.
Batch: machine learning inference serving on serverless platforms with adaptive batching
Ahsan Ali, Riccardo Pinciroli, Feng Yan, and Evgenia Smirni · 2020
Earlier work this paper cites.
Seuss: skip redundant paths to make serverless fast
James Cadden, Thomas Unger, Yara Awad, Han Dong, Orran Krieger, and Jonathan Appavoo · 2020
Earlier work this paper cites.
Catalyzer: Sub-millisecond startup for serverless computing with initialization-less booting
Dong Du, Tianyi Yu, Yubin Xia, Binyu Zang, Guanglu Yan, Chenggang Qin, Qixuan Wu, and Haibo Chen · 2020
Earlier work this paper cites.
Serving DNNs like Clockwork: Performance Predictability from the Bottom Up
Arpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao, Antoine Kaufmann, Ymir Vigfusson, and Jonathan Mace · 2020
Earlier work this paper cites.
{ \{ KungFu } \} : Making training in distributed machine learning adaptive
Luo Mai, Guo Li, Marcel Wagenländer, Konstantinos Fertakis, Andrei-Octavian Brabete, and Peter Pietzuch · 2020
Earlier work this paper cites.
Serverless in the wild: Characterizing and optimizing the serverless workload at a large cloud provider
Mohammad Shahrad, Rodrigo Fonseca, Inigo Goiri, Gohar Chaudhry, Paul Batum, Jason Cooke, Eduardo Laureano, Colby Tresness, Mark Russinovich, and Ricardo Bianchini · 2020
Earlier work this paper cites.
Faasm: Lightweight isolation for efficient stateful serverless computing
Simon Shillaker and Peter Pietzuch · 2020
Earlier work this paper cites.
Cloudburst: stateful functions-as-a-service
Vikram Sreekanti, Chenggang Wu, Xiayue Charles Lin, Johann Schleier-Smith, Joseph E. Gonzalez, Joseph M. Hellerstein, and Alexey Tumanov · 2020
Earlier work this paper cites.
Spotnik: Designing distributed machine learning for transient cloud resources
Marcel Wagenländer, Luo Mai, Guo Li, and Peter Pietzuch · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
When cloud storage meets { \{ rdma } \}
Yixiao Gao, Qiang Li, Lingbo Tang, Yongqing Xi, Pengcheng Zhang, Wenwen Peng, Bo Li, Yaohui Wu, Shaozong Liu, Lei Yan, et al · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Earlier work this paper cites.
Boki: Stateful serverless computing with shared logs
Zhipeng Jia and Emmett Witchel · 2021
Earlier work this paper cites.
SONIC: Application-aware data passing for chained serverless applications
Ashraf Mahgoub, Karthick Shankar, Subrata Mitra, Ana Klimovic, Somali Chaterji, and Saurabh Bagchi · 2021
Earlier work this paper cites.
Faa$t: A transparent auto-scaling cache for serverless applications
Francisco Romero, Gohar Irfan Chaudhry, Íñigo Goiri, Pragna Gopa, Paul Batum, Neeraja J. Yadwadkar, Rodrigo Fonseca, Christos Kozyrakis, and Ricardo Bianchini · 2021
Earlier work this paper cites.
INFaaS: Automated model-less inference serving
Francisco Romero, Qian Li, Neeraja J Yadwadkar, and Christos Kozyrakis · 2021
Earlier work this paper cites.
Benchmarking, analysis, and optimization of serverless function snapshots
Dmitrii Ustiugov, Plamen Petrov, Marios Kogias, Edouard Bugnion, and Boris Grot · 2021
Earlier work this paper cites.
FaaSNet: Scalable and fast provisioning of custom serverless container runtimes at alibaba cloud function compute
Ao Wang, Shuai Chang, Huangshi Tian, Hongqi Wang, Haoran Yang, Huiba Li, Rui Du, and Yue Cheng · 2021
Earlier work this paper cites.
Gillis: Serving large neural networks in serverless functions with automatic model partitioning
Minchen Yu, Zhifeng Jiang, Hok Chun Ng, Wei Wang, Ruichuan Chen, and Bo Li · 2021
Earlier work this paper cites.
Faster and cheaper serverless computing on harvested resources
Yanqi Zhang, Íñigo Goiri, Gohar Irfan Chaudhry, Rodrigo Fonseca, Sameh Elnikety, Christina Delimitrou, and Ricardo Bianchini · 2021
Earlier work this paper cites.
Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale
Reza Yazdani Aminabadi, Samyam Rajbhandari, Ammar Ahmad Awan, Cheng Li, Du Li, Elton Zheng, Olatunji Ruwase, Shaden Smith, Minjia Zhang, Jeff Rasley, and Yuxiong He · 2022
Earlier work this paper cites.
Faasnap: Faas made fast using snapshot-based vms
Lixiang Ao, George Porter, and Geoffrey M. Voelker · 2022
Earlier work this paper cites.
Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal Sharing
Seungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park, Youngjin Kwon, and Jaehyuk Huh · 2022
Earlier work this paper cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2022
Earlier work this paper cites.
Accelerate: Training and inference at scale made simple, efficient and adaptable
Sylvain Gugger, Lysandre Debut, Thomas Wolf, Philipp Schmid, Zachary Mueller, Sourab Mangrulkar, Marc Sun, and Benjamin Bossan · 2022
Cited alongside, same era.
Tetris: Memory-efficient serverless inference through tensor sharing
Jie Li, Laiping Zhao, Yanan Yang, Kunlin Zhan, and Keqiu Li · 2022
Cited alongside, same era.
Rund: A lightweight secure container runtime for high-density deployment and high-concurrency startup in serverless computing
Zijun Li, Jiagan Cheng, Quan Chen, Eryu Guan, Zizheng Bian, Yi Tao, Bin Zha, Qiang Wang, Weidong Han, and Minyi Guo · 2022
Cited alongside, same era.
ORION and the three rights: Sizing, bundling, and prewarming for serverless DAGs
Ashraf Mahgoub, Edgardo Barsallo Yi, Karthick Shankar, Sameh Elnikety, Somali Chaterji, and Saurabh Bagchi · 2022
Cited alongside, same era.
Peft: State-of-the-art parameter-efficient fine-tuning methods
Sourab Mangrulkar, Sylvain Gugger, Lysandre Debut, Younes Belkada, Sayak Paul, and Benjamin Bossan · 2022
Cited alongside, same era.
https://fireworks.ai/ , 2024
Generative AI tailored to you · 2024
Closest in time.
https://cohere.com/ , 2024
The leading enterprise AI platform · 2024
Closest in time.
https://replicate.com/ , 2024
Run AI with an API · 2024
Closest in time.
https://www.together.ai/products#inference , 2024
Serverless endpoints for leading open-source models · 2024
Closest in time.
https://www.databricks.com/ , 2024
Your data. your AI. your future · 2024
Closest in time.
Serverless inference - minimizing cold starts
Amazon · 2024
Closest in time.
Flexible I/O Tester, 2022
Jens Axboe · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Icebreaker: Warming serverless functions better with heterogeneity
Rohan Basu Roy, Tirthak Patel, and Devesh Tiwari · 2022
Cited alongside, same era.
Singularity: Planet-scale, preemptive and elastic scheduling of AI workloads, 2022
Dharma Shukla, Muthian Sivathanu, Srinidhi Viswanatha, Bhargav Gulavani, Rimma Nehme, Amey Agrawal, Chen Chen, Nipun Kwatra, Ramachandran Ramjee, Pankaj Sharma, Atul Katiyar, Vipul Modi, Vaibhav Sharma, Abhishek Singh, Shreshth Singhal, Kaustubh Welankar, Lu Xun, Ravi Anupindi, Karthik Elangovan, Hasibur Rahman, Zhou Lin, Rahul Seetharaman, Cheng Xu, Eddie Ailijiang, Suresh Krishnappa, and Mark Russinovich · 2022
Cited alongside, same era.
Infless: a native serverless system for low-latency, high-throughput inference
Yanan Yang, Laiping Zhao, Yiming Li, Huanyu Zhang, Jie Li, Mingyang Zhao, Xingzhen Chen, and Keqiu Li · 2022
Cited alongside, same era.
Orca: A distributed serving system for Transformer-Based generative models
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun · 2022
Cited alongside, same era.
OPT: Open Pre-trained Transformer Language Models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona T. Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer · 2022
Cited alongside, same era.
Palette load balancing: Locality hints for serverless functions
Mania Abdi, Samuel Ginzburg, Xiayue Charles Lin, Jose Faleiro, Gohar Irfan Chaudhry, Inigo Goiri, Ricardo Bianchini, Daniel S Berger, and Rodrigo Fonseca · 2023
Cited alongside, same era.
Falcon-40B: an open large language model with state-of-the-art performance
Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Merouane Debbah, Etienne Goffinet, Daniel Heslow, Julien Launay, Quentin Malartic, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo · 2023
Cited alongside, same era.
https://www.anyscale.com/blog/loading-llama-2-70b-20x-faster-with-anyscale-endpoints , 2023
Anyscale Technical Blog · 2024
Closest in time.
https://www.banana.dev/blog/turboboot , 2023
Banana.dev Technical Blog · 2024
Closest in time.
Reinventing search with a new AI-powered Microsoft Bing and Edge, your copilot for the web
Microsoft Official Blog · 2024
Closest in time.
https://github.com/features/copilot , 2023
GitHub Copilot · 2024
Closest in time.
LLM Inference Performance Engineering: Best Practices
DataBricks · 2024
Closest in time.
https://huggingface.co , 2023
HuggingFace · 2024
Closest in time.
Safetensors: ML Safer for All
HuggingFace · 2024
Closest in time.
Speed Comparison
HuggingFace · 2024
Closest in time.
Azure ML
Microsoft · 2024
Closest in time.
NVIDIA DGX H100
NVIDIA · 2024
Closest in time.
ChatGPT: Optimizing Language Models for Dialogue
OpenAI · 2024
Closest in time.
PyTorch Documentation: Saving and Loading Models
PyTorch · 2024
Closest in time.
torch.load
PyTorch · 2024
Closest in time.
API Reference, onnxruntime.backend.prepare
ONNX Runtime · 2024
Closest in time.
Déjàvu: Kv-cache streaming for fast, fault-tolerant generative LLM serving, 2024
Foteini Strati, Sara Mcallister, Amar Phanishayee, Jakub Tarnawski, and Ana Klimovic · 2024
Closest in time.
Better, faster, stronger
Mistral AI Team · 2024
Closest in time.
Introducing DBRX: A new state-of-the-art open LLM
The Mosaic Research Team · 2024
Closest in time.
Ray serve
The Ray Team · 2024
Closest in time.
TensorFlow Documentation: Using the Saved Model Format
TensorFlow · 2024
Closest in time.
tf.saved_model.load
TensorFlow · 2024
Closest in time.
Grok-1, 2024
xAI · 2024
Closest in time.