Fetching the paper…
Reading the bibliography…
Machine learning (ML) inference platforms are tasked with balancing two competing goals: ensuring high throughput given many requests, and delivering low-latency responses to support interactive applications.
DeeBERT: Dynamic Early Exiting for Accelerating BERT Inference
Ji Xin, Raphael Tang, Jaejun Lee, Yaoliang Yu, and Jimmy Lin. 2020 · 2004
Earlier work this paper cites.
Hill-climbing search
Bart Selman and Carla P Gomes. 2006 · 2006
Earlier work this paper cites.
Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, and B. Kingsbury. 2012 · 2012
Earlier work this paper cites.
Web data: Amazon reviews
2013 · 2013
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Adaptive neural networks for efficient inference. In International Conference on Machine Learning . PMLR, 527–536
Tolga Bolukbasi, Joseph Wang, Ofer Dekel, and Venkatesh Saligrama. 2017 · 2017
Earlier work this paper cites.
Clipper: A Low-Latency Online Prediction Serving System. In 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17) . USENIX Association, Boston, MA, 613–627
Daniel Crankshaw, Xin Wang, Guilio Zhou, Michael J. Franklin, Joseph E. Gonzalez, and Ion Stoica. 2017 · 2017
Earlier work this paper cites.
Modulating early visual processing by language
Harm de Vries, Florian Strub, Jérémie Mary, Hugo Larochelle, Olivier Pietquin, and Aaron Courville. 2017 · 2017
Earlier work this paper cites.
Swayam: Distributed Autoscaling to Meet SLAs of Machine Learning Inference Services with Resource Efficiency. In Proceedings of the 18th ACM/IFIP/USENIX Middleware Conference (Las Vegas, Nevada) (Middleware ’17) . Association for Computing Machinery, New York, NY, USA, 109–120
Arpan Gujarati, Sameh Elnikety, Yuxiong He, Kathryn S. McKinley, and Björn B. Brandenburg. 2017 · 2017
Earlier work this paper cites.
TensorFlow-Serving: Flexible, High-Performance ML Serving
Christopher Olston, Noah Fiedel, Kiril Gorovoy, Jeremiah Harmsen, Li Lao, Fangwei Li, Vinu Rajashekhar, Sukriti Ramesh, and Jordan Soyke. 2017 · 2017
Earlier work this paper cites.
Get to the point: Summarization with pointer-generator networks
Abigail See, Peter J Liu, and Christopher D Manning. 2017 · 2017
Earlier work this paper cites.
BranchyNet: Fast Inference via Early Exiting from Deep Neural Networks
Surat Teerapittayanon, Bradley McDanel, and H. T. Kung. 2017 · 2017
Earlier work this paper cites.
The Need for Mobile Speed: How Mobile Latency Impacts Publisher Revenue
Think with Google. 2017 · 2017
Earlier work this paper cites.
Neural Network Exchange Format (NNEF)
2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Applied Machine Learning at Facebook: A Datacenter Infrastructure Perspective. In 2018 IEEE International Symposium on High Performance Computer Architecture (HPCA) . 620–629
Kim Hazelwood, Sarah Bird, David Brooks, Soumith Chintala, Utku Diril, Dmytro Dzhulgakov, Mohamed Fawzy, Bill Jia, Yangqing Jia, Aditya Kalro, James Law, Kevin Lee, Jason Lu, Pieter Noordhuis, Misha Smelyanskiy, Liang Xiong, and Xiaodong Wang. 2018 · 2018
Earlier work this paper cites.
Focus: Querying Large Video Datasets with Low Latency and Low Cost. In Proceedings of the 13th USENIX Conference on Operating Systems Design and Implementation (Carlsbad, CA, USA) (OSDI’18) . USENIX Association, USA, 269–286
Kevin Hsieh, Ganesh Ananthanarayanan, Peter Bodik, Shivaram Venkataraman, Paramvir Bahl, Matthai Philipose, Phillip B. Gibbons, and Onur Mutlu. 2018 · 2018
Earlier work this paper cites.
SkipNet: Learning Dynamic Routing in Convolutional Networks
Xin Wang, Fisher Yu, Zi-Yi Dou, Trevor Darrell, and Joseph E. Gonzalez. 2018 · 2018
Earlier work this paper cites.
Dynamic Video Segmentation Network
Yu-Syuan Xu, Tsu-Jui Fu, Hsuan-Kung Yang, and Chun-Yi Lee. 2018 · 2018
Earlier work this paper cites.
Scaling video analytics on constrained edge nodes
Christopher Canel, Thomas Kim, Giulio Zhou, Conglong Li, Hyeontaek Lim, David G Andersen, Michael Kaminsky, and Subramanya Dulloor. 2019 · 2019
Earlier work this paper cites.
Analysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training Workloads. In 2019 USENIX Annual Technical Conference (USENIX ATC 19) . USENIX Association, Renton, WA, 947–960
Myeongjae Jeon, Shivaram Venkataraman, Amar Phanishayee, Junjie Qian, Wencong Xiao, and Fan Yang. 2019 · 2019
Earlier work this paper cites.
Shallow-deep networks: Understanding and mitigating network overthinking. In International conference on machine learning . PMLR, 3301–3310
Yigitcan Kaya, Sanghyun Hong, and Tudor Dumitras. 2019 · 2019
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 2019
Earlier work this paper cites.
Nexus: A GPU Cluster Engine for Accelerating DNN-Based Video Analysis. In Proceedings of the 27th ACM Symposium on Operating Systems Principles (Huntsville, Ontario, Canada) (SOSP ’19) . Association for Computing Machinery, New York, NY, USA, 322–337
Haichen Shen, Lequn Chen, Yuchen Jin, Liangyu Zhao, Bingyu Kong, Matthai Philipose, Arvind Krishnamurthy, and Ravi Sundaram. 2019 · 2019
Earlier work this paper cites.
Tactile Internet for Autonomous Vehicles: Latency and Reliability Analysis
Sudeep Tanwar, Sudhanshu Tyagi, Ishan Budhiraja, and Neeraj Kumar. 2019 · 2019
Earlier work this paper cites.
Machine Learning at Facebook: Understanding Inference at the Edge. In 2019 IEEE International Symposium on High Performance Computer Architecture (HPCA) . 331–344
Carole-Jean Wu, David Brooks, Kevin Chen, Douglas Chen, Sy Choudhury, Marat Dukhan, Kim Hazelwood, Eldad Isaac, Yangqing Jia, Bill Jia, Tommer Leyvand, Hao Lu, Yang Lu, Lin Qiao, Brandon Reagen, Joe Spisak, Fei Sun, Andrew Tulloch, Peter Vajda, Xiaodong Wang, Yanghan Wang, Bram Wasti, Yiming Wu, Ran Xian, Sungjoo Yoo, and Peizhao Zhang. 2019 · 2019
Earlier work this paper cites.
MArk: Exploiting Cloud Services for Cost-Effective, SLO-Aware Machine Learning Inference Serving. In 2019 USENIX Annual Technical Conference (USENIX ATC 19) . USENIX Association, Renton, WA, 1049–1062
Chengliang Zhang, Minchen Yu, Wei Wang, and Feng Yan. 2019 · 2019
Cited alongside, same era.
Preech: A System for Privacy-Preserving Speech Transcription. In 29th USENIX Security Symposium (USENIX Security 20) . USENIX Association, 2703–2720
Shimaa Ahmed, Amrita Roy Chowdhury, Kassem Fawaz, and Parmesh Ramanathan. 2020 · 2020
Cited alongside, same era.
InferLine: Latency-Aware Provisioning and Scaling for Prediction Serving Pipelines. In Proceedings of the 11th ACM Symposium on Cloud Computing (Virtual Event, USA) (SoCC ’20) . Association for Computing Machinery, New York, NY, USA, 477–491
Daniel Crankshaw, Gur-Eyal Sela, Xiangxi Mo, Corey Zumar, Ion Stoica, Joseph Gonzalez, and Alexey Tumanov. 2020 · 2020
Cited alongside, same era.
Serving DNNs like Clockwork: Performance Predictability from the Bottom Up. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20) . USENIX Association, 443–462
How Microsoft’s bet on Azure unlocked an AI revolution
2023 · 2023
Closest in time.
NVIDIA TensorRT: Programmable Inference Accelerator
2023 · 2023
Closest in time.
Boggart: Towards General-Purpose Acceleration of Retrospective Video Analytics. In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23) . USENIX Association, Boston, MA, 933–951
Neil Agarwal and Ravi Netravali. 2023 · 2023
Closest in time.
Fast and Robust Early-Exiting Framework for Autoregressive Language Models with Synchronized Parallel Decoding. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 5910–5924
Sangmin Bae, Jongwoo Ko, Hwanjun Song, and Se-Young Yun. 2023 · 2023
Closest in time.
Optimizing Dynamic Neural Networks with Brainstorm. In 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23) . USENIX Association, Boston, MA, 797–815
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Arpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao, Antoine Kaufmann, Ymir Vigfusson, and Jonathan Mace. 2020 · 2020
Cited alongside, same era.
FastBERT: a Self-distilling BERT with Adaptive Inference Time. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics, Online, 6035–6044
Weijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao, Haotang Deng, and Qi Ju. 2020 · 2020
Cited alongside, same era.
MLPerf: An Industry Standard Benchmark Suite for Machine Learning Performance
Peter Mattson, Vijay Janapa Reddi, Christine Cheng, Cody Coleman, Greg Diamos, David Kanter, Paulius Micikevicius, David Patterson, Guenther Schmuelling, Hanlin Tang, Gu-Yeon Wei, and Carole-Jean Wu. 2020 · 2020
Cited alongside, same era.
IMDb Movie Reviews Dataset
Aditya Pal, Abhilash Barigidad, and Abhijit Mustafi. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Cited alongside, same era.
The Right Tool for the Job: Matching Model and Instance Complexities. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . Association for Computational Linguistics, Online, 6640–6651
Roy Schwartz, Gabriel Stanovsky, Swabha Swayamdipta, Jesse Dodge, and Noah A. Smith. 2020 · 2020
Cited alongside, same era.
ODIN: Automated Drift Detection and Recovery in Video Analytics
Abhijit Suprem, Joy Arulraj, Calton Pu, and Joao Ferreira. 2020 · 2020
Cited alongside, same era.
Low Latency Speech Recognition using End-to-End Prefetching. In Interspeech 2020
Shuo yiin Chang, Bo Li, David Johannes Rybach, Wei Li, Yanzhang (Ryan) He, Tara N Sainath, and Trevor Deatrick Strohman. 2020 · 2020
Cited alongside, same era.
Wangchunshu Zhou, Canwen Xu, Tao Ge, Julian McAuley, Ke Xu, and Furu Wei. 2020 · 2020
Cited alongside, same era.
Weihao Cui, Zhenhua Han, Lingji Ouyang, Yichuan Wang, Ningxin Zheng, Lingxiao Ma, Yuqing Yang, Fan Yang, Jilong Xue, Lili Qiu, Lidong Zhou, Quan Chen, Haisheng Tan, and Minyi Guo. 2023 · 2023
Closest in time.
SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
Elias Frantar and Dan Alistarh. 2023 · 2023
Closest in time.
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. 2023 · 2023
Closest in time.
Amazon Found Every 100ms of Latency Cost them 1% in Sales
Gigaspaces. 2023 · 2023
Closest in time.
Pretrained Models
HuggingFace. 2023 · 2023
Closest in time.
RECL: Responsive Resource-Efficient Continuous Learning for Video Analytics. In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23) . USENIX Association, Boston, MA, 917–932
Mehrdad Khani, Ganesh Ananthanarayanan, Kevin Hsieh, Junchen Jiang, Ravi Netravali, Yuanchao Shu, Mohammad Alizadeh, and Victor Bahl. 2023 · 2023
Closest in time.
Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th Symposium on Operating Systems Principles . 611–626
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023 · 2023
Closest in time.
Fast Inference from Transformers via Speculative Decoding
Yaniv Leviathan, Matan Kalman, and Yossi Matias. 2023 · 2023
Closest in time.
{ \{ AlpaServe } \} : Statistical multiplexing with model parallelism for deep learning serving. In 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23) . 663–679
Zhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu, Ying Sheng, Xin Jin, Yanping Huang, Zhifeng Chen, Hao Zhang, Joseph E Gonzalez, et al · 2023
Closest in time.
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, Chuang Gan, and Song Han. 2023 · 2023
Closest in time.
LLM-Pruner: On the Structural Pruning of Large Language Models
Xinyin Ma, Gongfan Fang, and Xinchao Wang. 2023 · 2023
Closest in time.
ChatGPT continues to be one of the fastest-growing services ever
Jon Porter. 2023 · 2023
Closest in time.
Model Zoo
PyTorch. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Closest in time.
Tabi: An efficient multi-level inference system for large language models. In Proceedings of the Eighteenth European Conference on Computer Systems . 233–248
Yiding Wang, Kai Chen, Haisheng Tan, and Kun Guo. 2023 · 2023
Closest in time.
HuggingFace Pipelines
2024 · 2024
Closest in time.
NVIDIA Triton Inference Server
2024 · 2024
Closest in time.
The Yelp Reviews Dataset
2024 · 2024
Closest in time.
TorchServe
2024 · 2024
Closest in time.
Layer skip: Enabling early exit inference and self-speculative decoding
Mostafa Elhoushi, Akshat Shrivastava, Diana Liskovich, Basil Hosmer, Bram Wasti, Liangzhen Lai, Anas Mahmoud, Bilge Acun, Saurabh Agarwal, Ahmed Roman, et al · 2024
Closest in time.
Improving DNN Inference Throughput Using Practical, Per-Input Compute Adaptation. In Proceedings of the 30th ACM Symposium on Operating Systems Principles (Austin, TX, USA) (SOSP ’24)
Anand Iyer, Mingyu Guan, Yinwei Dai, Rui Pan, Swapnil Gandhi, and Ravi Netravali. 2024 · 2024
Closest in time.
SqueezeLLM: Dense-and-Sparse Quantization
Sehoon Kim, Coleman Hooper, Amir Gholami, Zhen Dong, Xiuyu Li, Sheng Shen, Michael W. Mahoney, and Kurt Keutzer. 2024 · 2024
Closest in time.