Fetching the paper…
Reading the bibliography…
As machine learning techniques are applied to a widening range of applications, high throughput machine learning (ML) inference servers have become critical for online service applications.
Bubble-Up: Increasing Utilization in Modern Warehouse Scale Computers via Sensible Co-locations
Jason Mars, Lingjia Tang, Robert Hundt, Kevin Skadron, and Mary Lou Soffa · 2011
Earlier work this paper cites.
Deep Learning Inference Platform
NVIDIA Corporation · 2013
Earlier work this paper cites.
Bubble-Flux: Precise Online QoS Management for Increased Utilization in Warehouse Scale Computers
Hailong Yang, Alex D. Breslow, Jason Mars, and Lingjia Tang · 2013
Earlier work this paper cites.
DjiNN and Tonic: DNN as a Service and Its Implications for Future Warehouse Scale Computers
Johann Hauswald, Yiping Kang, Michael A Laurenzano, Quan Chen, Cheng Li, Trevor Mudge, Ronald G Dreslinski, Jason Mars, and Lingjia Tang · 2015
Earlier work this paper cites.
Tensorflow: A System for Large-Scale Machine Learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2016
Earlier work this paper cites.
Baymax: QoS Awareness and Increased Utilization for Non-Preemptive Accelerators in Warehouse Scale Computers
Quan Chen, Hailong Yang, Jason Mars, and Lingjia Tang · 2016
Earlier work this paper cites.
AMD MULTIUSER GPU:HARDWARE-ENABLED GPU VIRTUALIZATION FOR A TRUE WORKSTATION EXPERIENCE
AMD Corporation · 2016
Earlier work this paper cites.
GeePS: Scalable Deep Learning on Distributed GPUs with a GPU-Specialized Parameter Server
Henggang Cui, Hao Zhang, Gregory R Ganger, Phillip B Gibbons, and Eric P Xing · 2016
Earlier work this paper cites.
Interference Management for Distributed Parallel Applications in Consolidated Clusters
Jaeung Han, Seungheun Jeon, Young ri Choi, and Jaehyuk Huh · 2016
Earlier work this paper cites.
Treadmill: Attributing the source of tail latency through precise load testing and statistical inference
Yunqi Zhang, David Meisner, Jason Mars, and Lingjia Tang · 2016
Earlier work this paper cites.
Prophet: Precise qos prediction on non-preemptive accelerators to improve utilization in warehouse-scale computers
Quan Chen, Hailong Yang, Minyi Guo, Ram Srivatsa Kannan, Jason Mars, and Lingjia Tang · 2017
Earlier work this paper cites.
Clipper: A Low-Latency Online Prediction Serving System
Daniel Crankshaw, Xin Wang, Guilio Zhou, Michael J Franklin, Joseph E Gonzalez, and Ion Stoica · 2017
Earlier work this paper cites.
In-datacenter performance analysis of a tensor processing unit
N. Jouppi et. al · 2017
Earlier work this paper cites.
Tensorflow-Serving: Flexible, High-Performance ML Serving
Christopher Olston, Noah Fiedel, Kiril Gorovoy, Jeremiah Harmsen, Li Lao, Fangwei Li, Vinu Rajashekhar, Sukriti Ramesh, and Jordan Soyke · 2017
Cited alongside, same era.
Poseidon: An Efficient Communication Architecture for Distributed Deep Learning on GPU Clusters
Hao Zhang, Zeyu Zheng, Shizhen Xu, Wei Dai, Qirong Ho, Xiaodan Liang, Zhiting Hu, Jinliang Wei, Pengtao Xie, and Eric P Xing · 2017
Cited alongside, same era.
TVM: An Automated End-to-End Optimizing Compiler for Deep Learning
Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy · 2018
Cited alongside, same era.
Applied machine learning at facebook: A datacenter infrastructure perspective
K. Hazelwood, S. Bird, D. Brooks, S. Chintala, U. Diril, D. Dzhulgakov, M. Fawzy, B. Jia, Y. Jia, A. Kalro, J. Law, K. Lee, J. Lu, P. Noordhuis, M. Smelyanskiy, L. Xiong, and X. Wang · 2018
Cited alongside, same era.
Dynamic Space-Time Scheduling for GPU Inference
Paras Jain, Xiangxi Mo, Ajay Jain, Harikaran Subbaraj, Rehan Durrani, Alexey Tumanov, Joseph Gonzalez, and Ion Stoica · 2018
A multi-neural network acceleration architecture
E. Baek, D. Kwon, and J. Kim · 2020
Later among the works it cites.
Prema: A predictive multi-task scheduling algorithm for preemptible neural processing units
Y. Choi and M. Rhu · 2020
Later among the works it cites.
Amazon SageMaker Developer Guide
Amazon Corporation · 2020
Later among the works it cites.
TensorRT Developer’s Guide
NVIDIA Corporation · 2020
Later among the works it cites.
Gslice: Controlled spatial sharing of gpus for a scalable inference platform
Aditya Dhakal, Sameer G Kulkarni, and K. K. Ramakrishnan · 2020
Later among the works it cites.
Planaria: Dynamic architecture fission for spatial multi-tenant acceleration of deep neural networks
S. Ghodrati, Byung Hoon Ahn, J. K. Kim, Sean Kinzer, Brahmendra Reddy Yatham, N. Alla, H. Sharma, Mohammad Alian, E. Ebrahimi, Nam Sung Kim, C. Young, and H. Esmaeilzadeh · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
PRETZEL: Opening the Black Box of Machine Learning Prediction Serving Systems
Yunseong Lee, Alberto Scolari, Byung-Gon Chun, Marco Domenico Santambrogio, Markus Weimer, and Matteo Interlandi · 2018
Cited alongside, same era.
Ray: A Distributed Framework for Emerging AI applications
Philipp Moritz, Robert Nishihara, Stephanie Wang, Alexey Tumanov, Richard Liaw, Eric Liang, Melih Elibol, Zongheng Yang, William Paul, Michael I. Jordan, and Ion Stoica · 2018
Cited alongside, same era.
Optimus: An Efficient Dynamic Resource Scheduler for Deep Learning Clusters
Yanghua Peng, Yixin Bao, Yangrui Chen, Chuan Wu, and Chuanxiong Guo · 2018
Cited alongside, same era.
Rafiki: Machine Learning as an Analytics Service System
Wei Wang, Sheng Wang, Jinyang Gao, Meihui Zhang, Gang Chen, Teck Ng, and Beng Ooi · 2018
Cited alongside, same era.
Gandiva: Introspective Cluster Scheduling for Deep Learning
Wencong Xiao, Romil Bhardwaj, Ramachandran Ramjee, Muthian Sivathanu, Nipun Kwatra, Zhenhua Han, Pratyush Patel, Xuan Peng, Hanyu Zhao, Quanlu Zhang, et al · 2018
Cited alongside, same era.
Multi-Process Service
NVIDIA Corporation · 2019
Cited alongside, same era.
Nexus: A GPU Cluster Engine for Accelerating DNN-Based Video Analysis
Haichen Shen, Lequn Chen, Yuchen Jin, Liangyu Zhao, Bingyu Kong, Matthai Philipose, Arvind Krishnamurthy, and Ravi Sundaram · 2019
Cited alongside, same era.
Later among the works it cites.
Serving dnns like clockwork: Performance predictability from the bottom up
Arpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao, Antoine Kaufmann, Ymir Vigfusson, and Jonathan Mace · 2020
Later among the works it cites.
Nimble: Lightweight and parallel GPU task scheduling for deep learning
Woosuk Kwon, Gyeong-In Yu, Eunji Jeong, and Byung-Gon Chun · 2020
Later among the works it cites.
Lazy batching: An sla-aware batching system for cloud machine learning inference
Yujeong Choi, Yunseong Kim, and Minsoo Rhu · 2021
Closest in time.
NVIDIA Multi-Instance GPU User Guide
NVIDIA Corporation · 2021
Closest in time.
Improving gpu multi-tenancy with page walk stealing
B Pratheek, Neha Jawalkar, and Arkaprava Basu · 2021
Closest in time.
Infaas: Automated model-less inference serving
Francisco Romero, Qian Li, Neeraja J. Yadwadkar, and Christos Kozyrakis · 2021
Closest in time.