Fetching the paper…
Reading the bibliography…
Recent years have witnessed an explosive growth of AI models.
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2021 · 2005
Earlier work this paper cites.
Achieving performance and availability guarantees with spot instances. In 2011 IEEE International Conference on High Performance Computing and Communications . IEEE, 296–303
Michele Mazzucco and Marlon Dumas. 2011 · 2011
Earlier work this paper cites.
SpotMPI: A Framework for Auction-Based HPC Computing Using Amazon Spot Instances. In Algorithms and Architectures for Parallel Processing: 11th International Conference, ICA300 2011, Melbourne, Australia, October 24-26, 2011, Proceedings, Part II 11 . Springer, 109–120
Moussa Taifi, Justin Y Shi, and Abdallah Khreishah. 2011 · 2011
Earlier work this paper cites.
Monetary cost-aware checkpointing and migration on Amazon cloud spot instances
Sangho Yi, Artur Andrzejak, and Derrick Kondo. 2011 · 2011
Earlier work this paper cites.
Optimal bidding in spot instance market. In 2012 Proceedings IEEE Infocom . IEEE, 190–198
Yang Song, Murtaza Zafer, and Kang-Won Lee. 2012 · 2012
Earlier work this paper cites.
A Reliable and Cost-Efficient Auto-Scaling System for Web Applications Using Heterogeneous Spot Instances
Chenhao Qu, Rodrigo N Calheiros, and Rajkumar Buyya. 2016 · 2016
Earlier work this paper cites.
Blending on-demand and spot instances to lower costs for in-memory storage. In IEEE INFOCOM 2016-The 35th Annual IEEE International Conference on Computer Communications . IEEE, 1–9
Zichen Xu, Christopher Stewart, Nan Deng, and Xiaorui Wang. 2016 · 2016
Earlier work this paper cites.
Proteus: agile ML elasticity through tiered reliability in dynamic resource markets. In Proceedings of the Twelfth European Conference on Computer Systems . 589–604
Aaron Harlap, Alexey Tumanov, Andrew Chung, Gregory R Ganger, and Phillip B Gibbons. 2017 · 2017
Earlier work this paper cites.
DeepSpotCloud: Leveraging Cross-Region GPU Spot Instances for Deep Learning. In 2017 IEEE 10th International Conference on Cloud Computing (CLOUD) . IEEE, 98–105
Kyungyong Lee and Myungjun Son. 2017 · 2017
Earlier work this paper cites.
Probabilistic guarantees of execution duration for Amazon spot instances. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis . 1–11
Rich Wolski, John Brevik, Ryan Chard, and Kyle Chard. 2017 · 2017
Earlier work this paper cites.
An efficient cost optimized scheduling for spot instances in heterogeneous cloud environment
Shridhar G Domanal and G Ram Mohana Reddy. 2018 · 2018
Earlier work this paper cites.
Tributary: spot-dancing for elastic services with latency SLOs. In 2018 USENIX Annual Technical Conference (USENIX ATC 18) . 1–14
Aaron Harlap, Andrew Chung, Alexey Tumanov, Gregory R Ganger, and Phillip B Gibbons. 2018 · 2018
Earlier work this paper cites.
Serving deep learning models in a serverless platform. In 2018 IEEE International conference on cloud engineering (IC2E) . IEEE, 257–262
Vatche Ishakian, Vinod Muthusamy, and Aleksander Slominski. 2018 · 2018
Earlier work this paper cites.
Optimizing the cost of executing mixed interactive and batch workloads on transient VMs
Pradeep Ambati and David Irwin. 2019 · 2019
Earlier work this paper cites.
Barista: Efficient and scalable serverless serving system for deep learning prediction services. In 2019 IEEE International Conference on Cloud Engineering (IC2E) . IEEE, 23–33
Anirban Bhattacharjee, Ajay Dev Chhokra, Zhuangwei Kang, Hongyang Sun, Aniruddha Gokhale, and Gabor Karsai. 2019 · 2019
Earlier work this paper cites.
MArk: Exploiting Cloud Services for Cost-Effective, SLO-Aware Machine Learning Inference Serving. In 2019 USENIX Annual Technical Conference (USENIX ATC 19) . 1049–1062
Chengliang Zhang, Minchen Yu, Wei Wang, and Feng Yan. 2019 · 2019
Earlier work this paper cites.
Serverless in the wild: Characterizing and optimizing the serverless workload at a large cloud provider. In 2020 USENIX annual technical conference (USENIX ATC 20) . 205–218
Mohammad Shahrad, Rodrigo Fonseca, Inigo Goiri, Gohar Chaudhry, Paul Batum, Jason Cooke, Eduardo Laureano, Colby Tresness, Mark Russinovich, and Ricardo Bianchini. 2020 · 2020
Earlier work this paper cites.
Spotnik: Designing Distributed Machine Learning for Transient Cloud Resources. In 12th USENIX Workshop on Hot Topics in Cloud Computing (HotCloud 20)
Marcel Wagenländer, Luo Mai, Guo Li, and Peter Pietzuch. 2020 · 2020
Earlier work this paper cites.
High-Resolution Image Synthesis with Latent Diffusion Models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2021 · 2021
Earlier work this paper cites.
Scheduling ML training on unreliable spot instances. In Proceedings of the 14th IEEE/ACM International Conference on Utility and Cloud Computing Companion . 1–8
Sheng Yang, Samir Khuller, Sunav Choudhary, Subrata Mitra, and Kanak Mahadik. 2021 · 2021
Earlier work this paper cites.
Varuna: Scalable, Low-cost Training of Massive Deep Learning Models. In Proceedings of the Seventeenth European Conference on Computer Systems . 472–487
Sanjith Athlur, Nitika Saran, Muthian Sivathanu, Ramachandran Ramjee, and Nipun Kwatra. 2022 · 2022
Earlier work this paper cites.
Cocktail: A Multidimensional Optimization for Model Serving in Cloud. In 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22) . 1041–1057
Jashwant Raj Gunasekaran, Cyan Subhra Mishra, Prashanth Thinakaran, Bikash Sharma, Mahmut Taylan Kandemir, and Chita R Das. 2022 · 2022
Earlier work this paper cites.
Dspy: Compiling declarative language model calls into self-improving pipelines
Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav Santhanam, Sri Vardhamanan, Saiful Haq, Ashutosh Sharma, Thomas T Joshi, Hanna Moazam, et al · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with paged attention. In Proceedings of the 29th Symposium on Operating Systems Principles . 611–626
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023 · 2023
Earlier work this paper cites.
AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving. In 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23) . 663–679
Zhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu, Ying Sheng, Xin Jin, Yanping Huang, Zhifeng Chen, Hao Zhang, Joseph E Gonzalez, et al · 2023
Earlier work this paper cites.
SpotServe: Serving Generative Large Language Models on Preemptible Instances
Xupeng Miao, Chunan Shi, Jiangfei Duan, Xiaoli Xi, Dahua Lin, Bin Cui, and Zhihao Jia. 2023 · 2023
Earlier work this paper cites.
The Inference Cost of Search Disruption–Large Language Model Cost Analysis
Dylan Patel and Afzal Ahmad. 2023 · 2023
Cited alongside, same era.
spotDNN: Provisioning Spot Instances for Predictable Distributed DNN Training in the Cloud. In 2023 IEEE/ACM 31st International Symposium on Quality of Service (IWQoS) . IEEE, 1–10
Ruitao Shang, Fei Xu, Zhuoyan Bai, Li Chen, Zhi Zhou, and Fangming Liu. 2023 · 2023
Cited alongside, same era.
Bamboo: Making Preemptible Instances Resilient for Affordable Training of Large DNNs. In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23) . 497–513
John Thorpe, Pengzhan Zhao, Jonathan Eyolfson, Yifan Qiao, Zhihao Jia, Minjia Zhang, Ravi Netravali, and Guoqing Harry Xu. 2023 · 2023
Cited alongside, same era.
Llama 2: Open Foundation and Fine-Tuned Chat Models
H. Touvron, L. Martin, K. Stone, and et al · 2023
Cited alongside, same era.
Loading Llama-2 70b 20x faster with Anyscale Endpoints
2024 · 2024
Closest in time.
Midjourney
2024 · 2024
Closest in time.
Navigating the High Cost of AI Compute
2024 · 2024
Closest in time.
NVIDIA Triton Inference Server
2024 · 2024
Closest in time.
OpenAI API
2024c · 2024
Closest in time.
Ray Serve: Scalable and Programmable Serving
2024 · 2024
Closest in time.
Running a GKE application on spot nodes with on-demand nodes as fallback
2024h · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Jeff Huang, Chuyue Sun, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E Gonzalez, et al · 2023
Cited alongside, same era.
Amazon EC2 Auto Scaling with EC2 Spot Instances
2024a · 2024
Cited alongside, same era.
Amazon EC2 Burstable Performance Instances
2024b · 2024
Cited alongside, same era.
Amazon SageMaker
2024c · 2024
Cited alongside, same era.
AWS Autoscaling Group
2024d · 2024
Cited alongside, same era.
AWS Lambda
2024e · 2024
Cited alongside, same era.
AWS Pricing
2024f · 2024
Cited alongside, same era.
AWS Spot Instance Advisor
2024g · 2024
Cited alongside, same era.
SpotServe: A Cost-Effective Spot Instance Serving Framework
2024 · 2024
Closest in time.
Text Generation Inference (TGI)
2024 · 2024
Closest in time.
Unofficial OpenAI Status
2024d · 2024
Closest in time.
Vertex AI
2024b · 2024
Closest in time.
Are more LLM calls all you need? Towards scaling laws of compound inference systems
Lingjiao Chen, Jared Quincy Davis, Boris Hanin, Peter Bailis, Ion Stoica, Matei Zaharia, and James Zou. 2024 · 2024
Closest in time.
Spot Virtual Machine Instances Documentation
Google Cloud. 2023 · 2024
Closest in time.
Agent AI: Surveying the Horizons of Multimodal Interaction
Zane Durante, Qiuyuan Huang, Naoki Wake, Ran Gong, Jae Sung Park, Bidipta Sarkar, Rohan Taori, Yusuke Noda, Demetri Terzopoulos, Yejin Choi, Katsushi Ikeuchi, Hoi Vo, Li Fei-Fei, and Jianfeng Gao. 2024 · 2024
Closest in time.
ServerlessLLM: Locality-Enhanced Serverless Inference for Large Language Models
Yao Fu, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete, Dmitrii Ustiugov, Yuvraj Patel, and Luo Mai. 2024 · 2024
Closest in time.
Use Azure Spot Virtual Machines
Microsoft. 2024 · 2024
Closest in time.
Behind the Scenes Scaling ChatGPT
Evan Morikawa. 2023 · 2024
Closest in time.
AI and Compute
OpenAI. 2024 · 2024
Closest in time.
LLM Routing: Bottleneck is Compute, Not the WAN
Mark Seery. 2024 · 2024
Closest in time.
Announcing Amazon EC2 Spot Instance Termination Notices
Amazon Web Services. 2015 · 2024
Closest in time.
Best Practices for Running Your Database on AWS Spot Instances
David Taylor. 2024 · 2024
Closest in time.
Can’t Be Late: Optimizing Spot Instance Savings under Deadlines. In 21st USENIX Symposium on Networked Systems Design and Implementation (NSDI 24) . 185–203
Zhanghao Wu, Wei-Lin Chiang, Ziming Mao, Zongheng Yang, Eric Friedman, Scott Shenker, and Ion Stoica. 2024 · 2024
Closest in time.
Anthropic: Claude 3.5 Sonnet
2025 · 2025
Closest in time.
Google Gemini: Supercharge your creativity and productivity
2025 · 2025
Closest in time.
Serverless Endpoints for leading open-source models
2025 · 2025
Closest in time.