Fetching the paper…
Reading the bibliography…
In the rapidly evolving landscape of artificial intelligence (AI), generative large language models (LLMs) stand at the forefront, revolutionizing how we interact with our data.
Speculative computation, parallelism, and functional programming
F Warren Burton. 1985 · 1985
Earlier work this paper cites.
SUMMA: Scalable universal matrix multiplication algorithm
Robert A Van De Geijn and Jerrell Watts. 1997 · 1997
Earlier work this paper cites.
Long short-term memory recurrent neural network architectures for large scale acoustic modeling. In Interspeech , Vol. 2014. 338–342
Hasim Sak, Andrew W Senior, and Françoise Beaufays. 2014 · 2014
Earlier work this paper cites.
Branchynet: Fast inference via early exiting from deep neural networks. In Proc. of ICPR 2016 . 2464–2469
Surat Teerapittayanon, Bradley McDanel, and Hsiang-Tsung Kung. 2016 · 2016
Earlier work this paper cites.
Language modeling with gated convolutional networks. In Proc. of ICML . 933–941
Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier. 2017 · 2017
Earlier work this paper cites.
The statistical recurrent unit. In Proc. of ICML . 2671–2680
Junier B Oliva, Barnabás Póczos, and Jeff Schneider. 2017 · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. In Proc. of ICLR 2017
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017 · 2017
Earlier work this paper cites.
Speeding up neural machine translation decoding by shrinking run-time vocabulary. In Proc. of ACL 2017 . 574–579
Xing Shi and Kevin Knight. 2017 · 2017
Earlier work this paper cites.
Attention is all you need. In Proc. of NeurIPS 2017
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Compiling machine learning programs via high-level tracing
Roy Frostig, Matthew James Johnson, and Chris Leary. 2018 · 2018
Earlier work this paper cites.
Non-autoregressive neural machine translation. In ICLR 2018
J Gu, J Bradbury, C Xiong, VOK Li, and R Socher. 2018 · 2018
Earlier work this paper cites.
Sharp Nearby, Fuzzy Far Away: How Neural Language Models Use Context. In Proc. of ACL 2018 . 284–294
Urvashi Khandelwal, He He, Peng Qi, and Dan Jurafsky. 2018 · 2018
Earlier work this paper cites.
Deterministic Non-Autoregressive Neural Sequence Modeling by Iterative Refinement. In Proc. of EMNLP 2018
Jason Lee, Elman Mansimov, and Kyunghyun Cho. 2018 · 2018
Earlier work this paper cites.
Online normalizer calculation for softmax
Maxim Milakov and Natalia Gimelshein. 2018 · 2018
Earlier work this paper cites.
Blockwise parallel decoding for deep autoregressive models
Mitchell Stern, Noam Shazeer, and Jakob Uszkoreit. 2018 · 2018
Earlier work this paper cites.
Demystifying parallel and distributed deep learning: An in-depth concurrency analysis
Tal Ben-Nun and Torsten Hoefler. 2019 · 2019
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
Adaptively Sparse Transformers. In Proc. of EMNLP-IJCNLP 2019
Gonçalo M Correia, Vlad Niculae, and André FT Martins. 2019 · 2019
Earlier work this paper cites.
Reducing Transformer Depth on Demand with Structured Dropout. In Proc. of ICLR 2019
Angela Fan, Edouard Grave, and Armand Joulin. 2019 · 2019
Earlier work this paper cites.
Mask-Predict: Parallel Decoding of Conditional Masked Language Models. In EMNLP-IJCNLP 2019 . 6112–6121
Marjan Ghazvininejad, Omer Levy, Yinhan Liu, and Luke Zettlemoyer. 2019 · 2019
Earlier work this paper cites.
TASO: optimizing deep learning computation with automatic generation of graph substitutions. In Proc. of SOSP 2019 . 47–62
Zhihao Jia, Oded Padon, James Thomas, Todd Warszawski, Matei Zaharia, and Alex Aiken. 2019a · 2019
Earlier work this paper cites.
Beyond Data and Model Parallelism for Deep Neural Networks
Zhihao Jia, Matei Zaharia, and Alex Aiken. 2019b · 2019
Earlier work this paper cites.
Reformer: The Efficient Transformer. In Proc. of ICLR 2019
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. 2019 · 2019
Earlier work this paper cites.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 2019
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 2019
Earlier work this paper cites.
Fast transformer decoding: One write-head is all you need
Noam Shazeer. 2019 · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019 · 2019
Earlier work this paper cites.
Patient Knowledge Distillation for BERT Model Compression. In Proc. of EMNLP-IJCNLP 2019 . 4323–4332
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019 · 2019
Earlier work this paper cites.
Triton: an intermediate language and compiler for tiled neural network computations. In Proceedings of the 3rd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages . 10–19
Philippe Tillet, Hsiang-Tsung Kung, and David Cox. 2019 · 2019
Earlier work this paper cites.
Sharing Attention Weights for Fast Transformer. In Proc. of IJCAI 2019 . 5292–5298
Tong Xiao, Yinqiao Li, Jingbo Zhu, Zhengtao Yu, and Tongran Liu. 2019 · 2019
Earlier work this paper cites.
MArk: Exploiting Cloud Services for Cost-Effective, SLO-Aware Machine Learning Inference Serving. In Proc. of USENIX ATC 2019 . 1049–1062
Chengliang Zhang, Minchen Yu, Wei Wang, and Feng Yan. 2019 · 2019
Earlier work this paper cites.
Batch: machine learning inference serving on serverless platforms with adaptive batching. In Proc. of SC 2020 . 1–15
Ahsan Ali, Riccardo Pinciroli, Feng Yan, and Evgenia Smirni. 2020 · 2020
Earlier work this paper cites.
PipeSwitch: Fast pipelined context switching for deep learning applications. In Proc. of OSDI 2020 . 499–514
Zhihao Bai, Zhen Zhang, Yibo Zhu, and Xin Jin. 2020 · 2020
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan. 2020 · 2020
Earlier work this paper cites.
Semi-autoregressive training improves mask-predict decoding
Marjan Ghazvininejad, Omer Levy, and Luke Zettlemoyer. 2020 · 2020
Earlier work this paper cites.
PoWER-BERT: Accelerating BERT inference via progressive word-vector elimination. In ICML 2020
Saurabh Goyal, Anamitra Roy Choudhury, Saurabh Raje, Venkatesan Chakaravarthy, Yogish Sabharwal, and Ashish Verma. 2020 · 2020
Earlier work this paper cites.
Gmat: Global memory augmentation for transformers
Ankit Gupta and Jonathan Berant. 2020 · 2020
Earlier work this paper cites.
TinyBERT: Distilling BERT for Natural Language Understanding. In Proc. of EMNLP Findings . 4163–4174
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020 · 2020
Earlier work this paper cites.
Deep Encoder, Shallow Decoder: Reevaluating Non-autoregressive Machine Translation. In Proc. of ICLR 2020
Jungo Kasai, Nikolaos Pappas, Hao Peng, James Cross, and Noah Smith. 2020 · 2020
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention. In Proc. of ICML 2020 . 5156–5165
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. 2020 · 2020
Earlier work this paper cites.
GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding. In Proc. of ICLR 2020
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2020 · 2020
Earlier work this paper cites.
Lei Li, Yankai Lin, Deli Chen, Shuhuai Ren, Peng Li, Jie Zhou, and Xu Sun. 2020 · 2020
Earlier work this paper cites.
FastBERT: a Self-distilling BERT with Adaptive Inference Time. In Proc. of ACL 2020 . 6035–6044
Weijie Liu, Peng Zhou, Zhiruo Wang, Zhe Zhao, Haotang Deng, and Qi Ju. 2020 · 2020
Earlier work this paper cites.
Scalable deep learning on distributed infrastructures: Challenges, techniques, and tools
Ruben Mayer and Hans-Arno Jacobsen. 2020 · 2020
Earlier work this paper cites.
Mlperf inference benchmark. In Proc. of ISCA 2020 . 446–459
Vijay Janapa Reddi, Christine Cheng, David Kanter, Peter Mattson, Guenther Schmuelling, Carole-Jean Wu, Brian Anderson, Maximilien Breughe, Mark Charlebois, William Chou, et al · 2020
Earlier work this paper cites.
Movement pruning: Adaptive sparsity by fine-tuning
Victor Sanh, Thomas Wolf, and Alexander Rush. 2020 · 2020
Earlier work this paper cites.
Sparse sinkhorn attention. In ICML 2020
Yi Tay, Dara Bahri, Liu Yang, Donald Metzler, and Da-Cheng Juan. 2020 · 2020
Earlier work this paper cites.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma. 2020a · 2020
Earlier work this paper cites.
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou. 2020b · 2020
Earlier work this paper cites.
LightSeq: A high performance inference library for transformers
Xiaohui Wang, Ying Xiong, Yang Wei, Mingxuan Wang, and Lei Li. 2020c · 2020
Earlier work this paper cites.
DeeBERT: Dynamic Early Exiting for Accelerating BERT Inference. In Proc. of ACL 2020 . 2246–2251
Ji Xin, Raphael Tang, Jaejun Lee, Yaoliang Yu, and Jimmy Lin. 2020 · 2020
Earlier work this paper cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al · 2020
Earlier work this paper cites.
Bert loses patience: Fast and robust inference with early exit
Wangchunshu Zhou, Canwen Xu, Tao Ge, Julian McAuley, Ke Xu, and Furu Wei. 2020 · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Et: re-thinking self-attention for transformer models on gpus. In Proc. of HPCA 2021 . 1–18
Shiyang Chen, Shaoyi Huang, Santosh Pandey, Bingbing Li, Guang R Gao, Long Zheng, Caiwen Ding, and Hang Liu. 2021a · 2021
Earlier work this paper cites.
Turbotransformers: an efficient gpu serving system for transformer models. In Proc. of PPoPP 2021 . 389–402
Jiarui Fang, Yang Yu, Chengduo Zhao, and Jie Zhou. 2021 · 2021
Earlier work this paper cites.
Efficiently Modeling Long Sequences with Structured State Spaces. In Proc. of ICLR 2021
Albert Gu, Karan Goel, and Christopher Re. 2021 · 2021
Earlier work this paper cites.
Fully Non-autoregressive Neural Machine Translation: Tricks of the Trade. In Proc. of ACL Findings . 120–133
Jiatao Gu and Xiang Kong. 2021 · 2021
Earlier work this paper cites.
Memory-efficient Transformers via Top-k Attention. In Proceedings of the Second Workshop on Simple and Efficient Natural Language Processing . 39–52
Ankit Gupta, Guy Dar, Shaya Goodman, David Ciprut, and Jonathan Berant. 2021 · 2021
Earlier work this paper cites.
Magic pyramid: Accelerating inference with early exiting and token pruning
Xuanli He, Iman Keivanloo, Yi Xu, Xiang He, Belinda Zeng, Santosh Rajagopalan, and Trishul Chilimbi. 2021 · 2021
Earlier work this paper cites.
Beyond Distillation: Task-level Mixture-of-Experts for Efficient Inference. In Proc. of EMNLP Findings . 3577–3599
Sneha Kudugunta, Yanping Huang, Ankur Bapna, Maxim Krikun, Dmitry Lepikhin, Minh-Thang Luong, and Orhan Firat. 2021 · 2021
Earlier work this paper cites.
An efficient transformer decoder with compressed sub-layers. In Proc. of AAAI 2021 . 13315–13323
Yanyang Li, Ye Lin, Tong Xiao, and Jingbo Zhu. 2021 · 2021
Earlier work this paper cites.
Accelerating sparse deep neural networks
Asit Mishra, Jorge Albericio Latorre, Jeff Pool, Darko Stosic, Dusan Stosic, Ganesh Venkatesh, Chong Yu, and Paulius Micikevicius. 2021 · 2021
Earlier work this paper cites.
Memory-efficient pipeline-parallel dnn training. In Proc. of ICML 2021 . 7937–7947
Deepak Narayanan, Amar Phanishayee, Kaiyu Shi, Xie Chen, and Matei Zaharia. 2021 · 2021
Earlier work this paper cites.
Evomoe: An evolutional mixture-of-experts training framework via dense-to-sparse gate
Xiaonan Nie, Xupeng Miao, Shijie Cao, Lingxiao Ma, Qibin Liu, Jilong Xue, Youshan Miao, Yi Liu, Zhi Yang, and Bin Cui. 2021 · 2021
Earlier work this paper cites.
Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation. In Proc. of ICLR 2021
Ofir Press, Noah Smith, and Mike Lewis. 2021 · 2021
Earlier work this paper cites.
Self-attention Does Not Need O( n 2 n^{2} ) Memory
Markus N Rabe and Charles Staats. 2021 · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation. In Proc. of ICML 2021 . 8821–8831
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021 · 2021
Earlier work this paper cites.
Hash layers for large sparse models
Stephen Roller, Sainbayar Sukhbaatar, Jason Weston, et al · 2021
Earlier work this paper cites.
Efficient content-based sparse attention with routing transformers
Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier. 2021 · 2021
Earlier work this paper cites.
Consistent Accelerated Inference via Confident Adaptive Transformers. In Proc. of EMNLP 2021 . 4962–4979
Tal Schuster, Adam Fisch, Tommi Jaakkola, and Regina Barzilay. 2021 · 2021
Earlier work this paper cites.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. 2021 · 2021
Earlier work this paper cites.
Mlp-mixer: An all-mlp architecture for vision
Ilya O Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, et al · 2021
Earlier work this paper cites.
TenTrans High-Performance Inference Toolkit for WMT2021 Efficiency Task. In Proceedings of the Sixth Conference on Machine Translation . 795–798
Kaixin Wu, Bojie Hu, and Qi Ju. 2021 · 2021
Earlier work this paper cites.
TR-BERT: Dynamic Token Reduction for Accelerating BERT Inference. In Proc. of NAACL 2021 . 5798–5809
Deming Ye, Yankai Lin, Yufei Huang, and Maosong Sun. 2021 · 2021
Earlier work this paper cites.
Shuangfei Zhai, Walter Talbott, Nitish Srivastava, Chen Huang, Hanlin Goh, Ruixiang Zhang, and Josh Susskind. 2021 · 2021
Earlier work this paper cites.
Informer: Beyond efficient transformer for long sequence time-series forecasting. In Proc. of AAAI 2021 . 11106–11115
Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang, Jianxin Li, Hui Xiong, and Wancai Zhang. 2021 · 2021
Earlier work this paper cites.
Deepspeed inference: Enabling efficient inference of transformer models at unprecedented scale
Reza Yazdani Aminabadi, Samyam Rajbhandari, Minjia Zhang, Ammar Ahmad Awan, Cheng Li, Du Li, Elton Zheng, Jeff Rasley, Shaden Smith, Olatunji Ruwase, et al · 2022
Earlier work this paper cites.
Improving language models by retrieving from trillions of tokens. In Proc. of ICML . 2206–2240
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al · 2022
Earlier work this paper cites.
Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W Cohen. 2022 · 2022
Earlier work this paper cites.
Accelerating transformer networks through recomposing softmax layers. In IISWC 2022 . 92–103
Jaewan Choi, Hailong Li, Byeongho Kim, Seunghwan Hwang, and Jung Ho Ahn. 2022 · 2022
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Earlier work this paper cites.
LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. 2022 · 2022
Earlier work this paper cites.
Glam: Efficient scaling of language models with mixture-of-experts. In Proc. of ICML 2022 . 5547–5569
Nan Du, Yanping Huang, Andrew M Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, et al · 2022
Earlier work this paper cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer. 2022 · 2022
Earlier work this paper cites.
The CoRa tensor compiler: Compilation for ragged tensors with minimal padding
Pratik Fegade, Tianqi Chen, Phillip Gibbons, and Todd Mowry. 2022 · 2022
Earlier work this paper cites.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. 2022a · 2022
Earlier work this paper cites.
OPTQ: Accurate quantization for generative pre-trained transformers. In Proc. of ICLR 2022
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. 2022b · 2022
Earlier work this paper cites.
Hungry Hungry Hippos: Towards Language Modeling with State Space Models. In Proc. of ICLR 2022
Daniel Y Fu, Tri Dao, Khaled Kamal Saab, Armin W Thomas, Atri Rudra, and Christopher Re. 2022 · 2022
Earlier work this paper cites.
Lossless acceleration for Seq2seq generation with aggressive decoding
Tao Ge, Heming Xia, Xin Sun, Si-Qing Chen, and Furu Wei. 2022 · 2022
Earlier work this paper cites.
A survey of quantization methods for efficient neural network inference
Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael W Mahoney, and Kurt Keutzer. 2022 · 2022
Earlier work this paper cites.
Cocktail: A multidimensional optimization for model serving in cloud. In Proc. of NSDI 2022 . 1041–1057
Jashwant Raj Gunasekaran, Cyan Subhra Mishra, Prashanth Thinakaran, Bikash Sharma, Mahmut Taylan Kandemir, and Chita R Das. 2022 · 2022
Earlier work this paper cites.
Compression of deep learning models for text: A survey
Manish Gupta and Puneet Agrawal. 2022 · 2022
Earlier work this paper cites.
Microsecond-scale preemption for concurrent GPU-accelerated DNN inferences. In Proc. of OSDI 2022 . 539–558
Mingcong Han, Hanze Zhang, Rong Chen, and Haibo Chen. 2022 · 2022
Earlier work this paper cites.
FasterMoE: modeling and optimizing training of large-scale dynamic pre-trained models. In Proc. of PPoPP 2022 . 120–134
Jiaao He, Jidong Zhai, Tiago Antunes, Haojie Wang, Fuwen Luo, Shangfeng Shi, and Qin Li. 2022 · 2022
Earlier work this paper cites.
MLIR-based code generation for GPU tensor cores. In Proceedings of the 31st ACM SIGPLAN International Conference on Compiler Construction . 117–128
Navdeep Katel, Vivek Khandelwal, and Uday Bondhugula. 2022 · 2022
Earlier work this paper cites.
Accelerating Inference for Pretrained Language Models by Unified Multi-Perspective Early Exiting. In Proc. of COLING . 4677–4686
Jun Kong, Jin Wang, Liang-Chih Yu, and Xuejie Zhang. 2022 · 2022
Earlier work this paper cites.
Long Range Language Modeling via Gated State Spaces. In Proc. of ICLR 2022
Harsh Mehta, Ankit Gupta, Ashok Cutkosky, and Behnam Neyshabur. 2022 · 2022
Earlier work this paper cites.
Adapler: Speeding up inference by adaptive length reduction
Ali Modarressi, Hosein Mohebbi, and Mohammad Taher Pilehvar. 2022 · 2022
Earlier work this paper cites.
LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models. In Proc. of ICLR 2022
Gunho Park, Minsub Kim, Sungjae Lee, Jeonghoon Kim, Beomseok Kwon, Se Jung Kwon, Byeongwook Kim, Youngjoo Lee, Dongsoo Lee, et al · 2022
Earlier work this paper cites.
Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale. In Proc. of ICML 2022 . 18332–18346
Samyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang, Reza Yazdani Aminabadi, Ammar Ahmad Awan, Jeff Rasley, and Yuxiong He. 2022 · 2022
Earlier work this paper cites.
Apache TVM Unity: a vision for the ML software and hardware ecosystem
Adrian Sampson, Tianqi Chen, and Jared Roesch. 2022 · 2022
Earlier work this paper cites.
A simple hash-based early exiting approach for language understanding and generation
Tianxiang Sun, Xiangyang Liu, Wei Zhu, Zhichao Geng, Lingling Wu, Yilong He, Yuan Ni, Guotong Xie, Xuanjing Huang, and Xipeng Qiu. 2022 · 2022
Earlier work this paper cites.
Unity: Accelerating DNN training through joint optimization of algebraic transformations and parallelization. In Proc. of OSDI 2022 . 267–284
Colin Unger, Zhihao Jia, Wei Wu, Sina Lin, Mandeep Baines, Carlos Efrain Quintero Narvaez, Vinay Ramakrishnaiah, Nirmal Prajapati, Pat McCormick, Jamaludin Mohd-Yusof, et al · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
MLaaS in the wild: Workload analysis and scheduling in Large-Scale heterogeneous GPU clusters. In Proc. of NSDI 2022 . 945–960
Qizhen Weng, Wencong Xiao, Yinghao Yu, Wei Wang, Cheng Wang, Jian He, Yong Li, Liping Zhang, Wei Lin, and Yu Ding. 2022 · 2022
Earlier work this paper cites.
Bloom: A 176b-parameter open-access multilingual language model
BigScience Workshop, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, et al · 2022
Earlier work this paper cites.
Speeding up Transformer Decoding via an Attention Refinement Network. In Proc. of COLING 2022 . 5109–5118
Kaixin Wu, Yue Zhang, Bojie Hu, and Tong Zhang. 2022 · 2022
Earlier work this paper cites.
Smoothquant: Accurate and efficient post-training quantization for large language models
Guangxuan Xiao, Ji Lin, Mickael Seznec, Julien Demouth, and Song Han. 2022 · 2022
Earlier work this paper cites.
Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Zhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu, Conglong Li, and Yuxiong He. 2022 · 2022
Earlier work this paper cites.
Orca: A Distributed Serving System for Transformer-Based Generative Models. In Proc. of OSDI 2022 . 521–538
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun. 2022a · 2022
Earlier work this paper cites.
Metaformer is actually what you need for vision. In Proc. of CVPR 2022 . 10819–10829
Weihao Yu, Mi Luo, Pan Zhou, Chenyang Si, Yichen Zhou, Xinchao Wang, Jiashi Feng, and Shuicheng Yan. 2022b · 2022
Earlier work this paper cites.
Glm-130b: An open bilingual pre-trained model
Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, et al · 2022
Earlier work this paper cites.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Earlier work this paper cites.
Alpa: Automating inter-and Intra-Operator parallelism for distributed deep learning. In Proc. of OSDI 2022 . 559–578
Lianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang, Zhifeng Chen, Yanping Huang, Yida Wang, Yuanzhong Xu, Danyang Zhuo, Eric P Xing, et al · 2022
Earlier work this paper cites.
Transpim: A memory-based acceleration via software-hardware co-design for transformer. In Proc. of HPCA 2022 . 1071–1085
Minxuan Zhou, Weihong Xu, Jaeyoung Kang, and Tajana Rosing. 2022c · 2022
Earlier work this paper cites.
Mixture-of-experts with expert choice routing
Yanqi Zhou, Tao Lei, Hanxiao Liu, Nan Du, Yanping Huang, Vincent Zhao, Andrew M Dai, Quoc V Le, James Laudon, et al · 2022
Earlier work this paper cites.
PetS: A Unified Framework for Parameter-Efficient Transformers Serving. In Proc. of USENIX ATC 2022 . 489–504
Zhe Zhou, Xuechao Wei, Jiejing Zhang, and Guangyu Sun. 2022b · 2022
Earlier work this paper cites.
NVIDIA Effective Transformer
2020 · 2023
Earlier work this paper cites.
NVIDIA FasterTransformer
2021 · 2023
Earlier work this paper cites.
DeepSpeed Inference
2022 · 2023
Earlier work this paper cites.
NVIDIA H100 Tensor Core GPU Architecture
2022 · 2023
Earlier work this paper cites.
AnyScale LLMPerf leaderboard
2023 · 2023
Cited alongside, same era.
AWS Inferentia
2023 · 2023
Cited alongside, same era.
ChatGLM2-6B
2023 · 2023
Cited alongside, same era.
CTranslate2
2023 · 2023
Cited alongside, same era.
DeepSpeed-FastGen
2023a · 2023
Cited alongside, same era.
DeepSpeed-Inference v.s. ZeRO-Inference
2023 · 2023
Cited alongside, same era.
DeepSpeed-MII
2023b · 2023
Cited alongside, same era.
FlexFlow-Serve
2023a · 2023
Efficient LLM Inference on CPUs
Haihao Shen, Hanwen Chang, Bo Dong, Yu Luo, and Hengyu Meng. 2023 · 2023
Closest in time.
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
Ying Sheng, Shiyi Cao, Dacheng Li, Coleman Hooper, Nicholas Lee, Shuo Yang, Christopher Chou, Banghua Zhu, Lianmin Zheng, Kurt Keutzer, et al · 2023
Closest in time.
FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU. In Proc. of ICML 2023 . 31094–31116
Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Ré, Ion Stoica, and Ce Zhang. 2023b · 2023
Closest in time.
Welder: Scheduling Deep Learning Memory Access via Tile-graph. In Proc. of OSDI 2023 . 701–718
Yining Shi, Zhi Yang, Jilong Xue, Lingxiao Ma, Yuqing Xia, Ziming Miao, Yuxiao Guo, Fan Yang, and Lidong Zhou. 2023 · 2023
Closest in time.
PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
gpt-fast
2023 · 2023
Cited alongside, same era.
Graphcore
2023 · 2023
Cited alongside, same era.
Graphcore PopTransformer
2023 · 2023
Cited alongside, same era.
Huggingface Text Generation Inference
2023 · 2023
Cited alongside, same era.
Yixin Song, Zeyu Mi, Haotong Xie, and Haibo Chen. 2023 · 2023
Closest in time.
Accelerating llm inference with staged speculative decoding
Benjamin Spector and Chris Re. 2023 · 2023
Closest in time.
A Simple and Effective Pruning Approach for Large Language Models
Mingjie Sun, Zhuang Liu, Anna Bair, and J Zico Kolter. 2023b · 2023
Closest in time.
Retentive Network: A Successor to Transformer for Large Language Models
Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang, and Furu Wei. 2023a · 2023
Closest in time.
FusionAI: Decentralized Training and Deploying LLMs with Massive Consumer-Level GPUs
Zhenheng Tang, Yuxin Wang, Xin He, Longteng Zhang, Xinglin Pan, Qiang Wang, Rongfei Zeng, Kaiyong Zhao, Shaohuai Shi, Bingsheng He, et al · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023 · 2023
Closest in time.
Efficient Transformers: A Survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. 2023 · 2023
Closest in time.
MLC-LLM
MLC team. 2023 · 2023
Closest in time.
AutoML in the Age of Large Language Models: Current Challenges, Future Opportunities and Risks
Alexander Tornede, Difan Deng, Theresa Eimer, Joseph Giovanelli, Aditya Mohan, Tim Ruhkopf, Sarah Segel, Daphne Theodorakopoulos, Tanja Tornede, Henning Wachsmuth, et al · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Closest in time.
Efficient methods for natural language processing: A survey
Marcos Treviso, Ji-Ung Lee, Tianchu Ji, Betty van Aken, Qingqing Cao, Manuel R Ciosici, Michael Hassid, Kenneth Heafield, Sara Hooker, Colin Raffel, et al · 2023
Closest in time.
Flash-Decoding for long-context inference
Francisco Massa Grigory Sizov Tri Dao, Daniel Haziza. 2023 · 2023
Closest in time.
Mini-GPTs: Efficient Large Language Models through Contextual Pruning
Tim Valicenti, Justice Vidal, and Ritik Patnaik. 2023 · 2023
Closest in time.
Tabi: An Efficient Multi-Level Inference System for Large Language Models. In Proc. of EuroSys 2023 . 233–248
Yiding Wang, Kai Chen, Haisheng Tan, and Kun Guo. 2023 · 2023
Closest in time.
Fast Distributed Inference Serving for Large Language Models
Bingyang Wu, Yinmin Zhong, Zili Zhang, Gang Huang, Xuanzhe Liu, and Xin Jin. 2023c · 2023
Closest in time.
PyTorch 2.0: The Journey to Bringing Compiler Technologies to the Core of PyTorch (Keynote). In Proceedings of the 21st ACM/IEEE International Symposium on Code Generation and Optimization . 1–1
Peng Wu. 2023 · 2023
Closest in time.
Autogen: Enabling next-gen llm applications via multi-agent conversation framework
Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Shaokun Zhang, Erkang Zhu, Beibin Li, Li Jiang, Xiaoyun Zhang, and Chi Wang. 2023a · 2023
Closest in time.
Xiaoxia Wu, Cheng Li, Reza Yazdani Aminabadi, Zhewei Yao, and Yuxiong He. 2023b · 2023
Closest in time.
Haojun Xia, Zhen Zheng, Yuchao Li, Donglin Zhuang, Zhongzhu Zhou, Xiafei Qiu, Yong Li, Wei Lin, and Shuaiwen Leon Song. 2023 · 2023
Closest in time.
Efficient Streaming Language Models with Attention Sinks
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. 2023a · 2023
Closest in time.
A survey on non-autoregressive generation for neural machine translation and beyond
Yisheng Xiao, Lijun Wu, Junliang Guo, Juntao Li, Min Zhang, Tao Qin, and Tie-yan Liu. 2023b · 2023
Closest in time.
Wizardlm: Empowering large language models to follow complex instructions
Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang. 2023c · 2023
Closest in time.
LLMCad: Fast and Scalable On-device Large Language Model Inference
Daliang Xu, Wangsong Yin, Xin Jin, Ying Zhang, Shiyun Wei, Mengwei Xu, and Xuanzhe Liu. 2023d · 2023
Closest in time.
Retrieval meets Long Context Large Language Models
Peng Xu, Wei Ping, Xianchao Wu, Lawrence McAfee, Chen Zhu, Zihan Liu, Sandeep Subramanian, Evelina Bakhturina, Mohammad Shoeybi, and Bryan Catanzaro. 2023b · 2023
Closest in time.
Zhaozhuo Xu, Zirui Liu, Beidi Chen, Yuxin Tang, Jue Wang, Kaixiong Zhou, Xia Hu, and Anshumali Shrivastava. 2023a · 2023
Closest in time.
Baichuan 2: Open large-scale language models
Aiyuan Yang, Bin Xiao, Bingning Wang, Borong Zhang, Chao Yin, Chenxu Lv, Da Pan, Dian Wang, Dong Yan, Fan Yang, et al · 2023
Closest in time.
Inference with reference: Lossless acceleration of large language models
Nan Yang, Tao Ge, Liang Wang, Binxing Jiao, Daxin Jiang, Linjun Yang, Rangan Majumder, and Furu Wei. 2023a · 2023
Closest in time.
Predictive Pipelined Decoding: A Compute-Latency Trade-off for Exact LLM Decoding
Seongjun Yang, Gibbeum Lee, Jaewoong Cho, Dimitris Papailiopoulos, and Kangwook Lee. 2023b · 2023
Closest in time.
A comprehensive study on post-training quantization for large language models
Zhewei Yao, Cheng Li, Xiaoxia Wu, Stephen Youn, and Yuxiong He. 2023 · 2023
Closest in time.
SparseTIR: Composable abstractions for sparse compilation in deep learning. In Proc. of ASPLOS 2023 . 660–678
Zihao Ye, Ruihang Lai, Junru Shao, Tianqi Chen, and Luis Ceze. 2023 · 2023
Closest in time.
A Scalable GPT-2 Inference Hardware Architecture on FPGA. In Proc. of IJCNN 2023 . 1–8
Anil Yemme and Shayan Srinivasa Garani. 2023 · 2023
Closest in time.
Stateful large language model serving with pensieve
Lingfan Yu, Jinkun Lin, and Jinyang Li. 2023 · 2023
Closest in time.
RPTQ: Reorder-based Post-training Quantization for Large Language Models
Zhihang Yuan, Lin Niu, Jiawei Liu, Wenyu Liu, Xinggang Wang, Yuzhang Shang, Guangyu Sun, Qiang Wu, Jiaxiang Wu, and Bingzhe Wu. 2023 · 2023
Closest in time.
Learning to Skip for Language Modeling
Dewen Zeng, Nan Du, Tao Wang, Yuanzhong Xu, Tao Lei, Zhifeng Chen, and Claire Cui. 2023 · 2023
Closest in time.
Bytetransformer: A high-performance transformer boosted for variable-length inputs. In IPDPS 2023 . 344–355
Yujia Zhai, Chengquan Jiang, Leyuan Wang, Xiaoying Jia, Shang Zhang, Zizhong Chen, Xin Liu, and Yibo Zhu. 2023 · 2023
Closest in time.
DePA: Improving Non-autoregressive Translation with Dependency-Aware Decoder. In Proc. of IWSLT 2023) . 478–490
Jiaao Zhan, Qian Chen, Boxing Chen, Wen Wang, Yu Bai, and Yang Gao. 2023 · 2023
Closest in time.
Longteng Zhang, Xiang Liu, Zeyu Li, Xinglin Pan, Peijie Dong, Ruibo Fan, Rui Guo, Xin Wang, Qiong Luo, Shaohuai Shi, et al · 2023
Closest in time.
Mengke Zhang, Tianxing He, Tianle Wang, Fatemehsadat Mireshghallah, Binyi Chen, Hao Wang, and Yulia Tsvetkov. 2023a · 2023
Closest in time.
H _ 2 \_2 O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models
Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, et al · 2023
Closest in time.
EINNET: Optimizing Tensor Programs with Derivation-Based Transformations. In Proc. of OSDI 2023 . 739–755
Liyan Zheng, Haojie Wang, Jidong Zhai, Muyan Hu, Zixuan Ma, Tuowei Wang, Shuhong Huang, Xupeng Miao, Shizhi Tang, Kezhao Huang, et al · 2023
Closest in time.
Efficiently Programming Large Language Models using SGLang
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Jeff Huang, Chuyue Sun, Cody_Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E Gonzalez, et al · 2023
Closest in time.
PIT: Optimization of Dynamic Sparse Deep Learning Models via Permutation Invariant Transformation. In Proc. of SOSP 2023 . 331–347
Ningxin Zheng, Huiqiang Jiang, Quanlu Zhang, Zhenhua Han, Lingxiao Ma, Yuqing Yang, Fan Yang, Chengruidong Zhang, Lili Qiu, Mao Yang, et al · 2023
Closest in time.
DistillSpec: Improving Speculative Decoding via Knowledge Distillation
Yongchao Zhou, Kaifeng Lyu, Ankit Singh Rawat, Aditya Krishna Menon, Afshin Rostamizadeh, Sanjiv Kumar, Jean-François Kagy, and Rishabh Agarwal. 2023 · 2023
Closest in time.
On Optimal Caching and Model Multiplexing for Large Model Inference
Banghua Zhu, Ying Sheng, Lianmin Zheng, Clark Barrett, Michael I Jordan, and Jiantao Jiao. 2023c · 2023
Closest in time.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023a · 2023
Closest in time.
A survey on model compression for large language models
Xunyu Zhu, Jian Li, Yong Liu, Can Ma, and Weiping Wang. 2023b · 2023
Closest in time.
Falcon LLM: A New Frontier in Natural Language Processing
Yoshua X ZXhang, Yann M Haxo, and Ying X Mat. 2023 · 2023
Closest in time.
Taming Throughput-Latency tradeoff in LLM inference with Sarathi-Serve. In Proc. of OSDI 2024 . 117–134
Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav Gulavani, Alexey Tumanov, and Ramachandran Ramjee. 2024 · 2024
Closest in time.
Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding. In COLM 2024
Zachary Ankner, Rishab Parthasarathy, Aniruddha Nrusimha, Christopher Rinard, Jonathan Ragan-Kelley, and William Brandon. 2024 · 2024
Closest in time.
LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding. In Proc. of ACL 2024
Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, et al · 2024
Closest in time.
Sequoia: Scalable and robust speculative decoding
Zhuoming Chen, Avner May, Ruslan Svirschevski, Yu-Hsun Huang, Max Ryabinin, Zhihao Jia, and Beidi Chen. 2024a · 2024
Closest in time.
Magicpig: Lsh sampling for efficient llm generation. In Proc. of ICLR 2024
Zhuoming Chen, Ranajoy Sadhukhan, Zihao Ye, Yang Zhou, Jianyu Zhang, Niklas Nolte, Yuandong Tian, Matthijs Douze, Leon Bottou, Zhihao Jia, et al · 2024
Closest in time.
LLM-Inference-Bench: Inference Benchmarking of Large Language Models on AI Accelerators. In Workshops of SC 2024 . 1362–1379
Krishna Teja Chitty-Venkata, Siddhisanket Raskar, Bharat Kale, Farah Ferdaus, Aditya Tanikanti, Ken Raffenetti, Valerie Taylor, Murali Emani, and Venkatram Vishwanath. 2024 · 2024
Closest in time.
Speculative diffusion decoding: Accelerating language generation through diffusion
Jacob K Christopher, Brian R Bartoldson, Tal Ben-Nun, Michael Cardei, Bhavya Kailkhura, and Ferdinando Fioretto. 2024 · 2024
Closest in time.
Flex Attention: A Programming Model for Generating Optimized Attention Kernels
Juechu Dong, Boyuan Feng, Driss Guessous, Yanbo Liang, and Horace He. 2024a · 2024
Closest in time.
Xgrammar: Flexible and efficient structured generation engine for large language models
Yixin Dong, Charlie F Ruan, Yaxing Cai, Ruihang Lai, Ziyi Xu, Yilong Zhao, and Tianqi Chen. 2024b · 2024
Closest in time.
Break the sequential dependency of LLM inference using LOOKAHEAD DECODING. In Proc. of ICML 2024 . 14060–14079
Yichao Fu, Peter Bailis, Ion Stoica, and Hao Zhang. 2024a · 2024
Closest in time.
Efficiently Serving LLM Reasoning Programs with Certaindex
Yichao Fu, Junda Chen, Siqi Zhu, Zheyu Fu, Zhongdongming Dai, Aurick Qiao, and Hao Zhang. 2024b · 2024
Closest in time.
ServerlessLLM:Low-Latency serverless inference for large language models. In Proc. of OSDI 2024 . 135–153
Yao Fu, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete, Dmitrii Ustiugov, Yuvraj Patel, and Luo Mai. 2024c · 2024
Closest in time.
Cost-Efficient large language model serving for multi-turn conversations with CachedAttention. In Proc. of USENIX ATC 2024 . 111–126
Bin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang, Djordje Jevdjic, Junbo Deng, Xingkun Yang, Zhou Yu, and Pengfei Zuo. 2024 · 2024
Closest in time.
Mamba: Linear-Time Sequence Modeling with Selective State Spaces. In COLM 2024
Albert Gu and Tri Dao. 2023 · 2024
Closest in time.
MiniLLM: Knowledge Distillation of Large Language Models. In Proc. of ICLR 2024
Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang. 2024 · 2024
Closest in time.
Inference without interference: Disaggregate llm inference for mixed downstream workloads
Cunchen Hu, Heyang Huang, Liangliang Xu, Xusheng Chen, Jiang Xu, Shuang Chen, Hao Feng, Chenxi Wang, Sa Wang, Yungang Bao, et al · 2024
Closest in time.
LANTERN: Accelerating Visual Autoregressive Models with Relaxed Speculative Decoding. In Proc. of ICLR 2024
Doohyuk Jang, Sihwan Park, June Yong Yang, Yeonsung Jung, Jihun Yun, Souvik Kundu, Sung-Yub Kim, and Eunho Yang. 2024 · 2024
Closest in time.
LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression. In ACL 2024
Huiqiang Jiang, Qianhui Wu, Xufang Luo, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. 2024 · 2024
Closest in time.
A System for Microserving of LLMs
Hongyi Jin, Ruihang Lai, Charlie F Ruan, Yingcheng Wang, Todd C Mowry, Xupeng Miao, Zhihao Jia, and Tianqi Chen. 2024 · 2024
Closest in time.
InfiniGen: Efficient generative inference of large language models with dynamic KV cache management. In Proc. of OSDI 2024 . 155–172
Wonbeom Lee, Jungi Lee, Junghwan Seo, and Jaewoong Sim. 2024 · 2024
Closest in time.
EAGLE: speculative sampling requires rethinking feature uncertainty. In Proc. of ICML 2024 . 28935–28948
Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang. 2024 · 2024
Closest in time.
Parrot: Efficient Serving of LLM-based Applications with Semantic Variable. In Proc. of OSDI 2024 . 929–945
Chaofan Lin, Zhenhua Han, Chengruidong Zhang, Yuqing Yang, Fan Yang, Chen Chen, and Lili Qiu. 2024a · 2024
Closest in time.
Awq: Activation-aware weight quantization for on-device llm compression and acceleration
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. 2024b · 2024
Closest in time.
Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model
Aixin Liu, Bei Feng, Bin Wang, Bingxuan Wang, Bo Liu, Chenggang Zhao, Chengqi Dengr, Chong Ruan, Damai Dai, Daya Guo, et al · 2024
Closest in time.
Kangaroo: Lossless self-speculative decoding for accelerating LLMs via double early exiting
Fangcheng Liu, Yehui Tang, Zhenhua Liu, Yunsheng Ni, Duyu Tang, Kai Han, and Yunhe Wang. 2024d · 2024
Closest in time.
Andes: Defining and enhancing quality-of-experience in llm-based text streaming services
Jiachen Liu, Jae-Won Chung, Zhiyu Wu, Fan Lai, Myungjin Lee, and Mosharaf Chowdhury. 2024a · 2024
Closest in time.
CacheGen: Fast Context Loading for Language Model Applications. In SIGCOMM 2024
Yuhan Liu, Hanchen Li, Kuntai Du, Jiayi Yao, Yihua Cheng, Yuyang Huang, Shan Lu, Michael Maire, Henry Hoffmann, Ari Holtzman, et al · 2024
Closest in time.
Skyserve: Serving ai models across regions and clouds with spot instances
Ziming Mao, Tian Xia, Zhanghao Wu, Wei-Lin Chiang, Tyler Griggs, Romil Bhardwaj, Zongheng Yang, Scott Shenker, and Ion Stoica. 2024 · 2024
Closest in time.
Realhf: Optimized rlhf training for large language models through parameter reallocation
Zhiyu Mei, Wei Fu, Kaiwei Li, Guangju Wang, Huanchen Zhang, and Yi Wu. 2024 · 2024
Closest in time.
Demystifying data management for large language models. In Companion of the 2024 International Conference on Management of Data . 547–555
Xupeng Miao, Zhihao Jia, and Bin Cui. 2024a · 2024
Closest in time.
FlexLLM: A System for Co-Serving Large Language Model Inference and Parameter-Efficient Finetuning
Xupeng Miao, Gabriele Oliaro, Xinhao Cheng, Mengdi Wu, Colin Unger, and Zhihao Jia. 2024b · 2024
Closest in time.
Specinfer: Accelerating large language model serving with tree-based speculative inference and verification. In Proc. of ASPLOS 2024 . 932–949
Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, et al · 2024
Closest in time.
SpotServe: Serving Generative Large Language Models on Preemptible Instances. In Proc. of ASPLOS 2024 . 1112–1127
Xupeng Miao, Chunan Shi, Jiangfei Duan, Xiaoli Xi, Dahua Lin, Bin Cui, and Zhihao Jia. 2024d · 2024
Closest in time.
Exegpt: Constraint-aware resource scheduling for llm inference. In Proc. of ASPLOS 2024 . 369–384
Hyungjun Oh, Kihong Kim, Jaemin Kim, Sungkyun Kim, Junyeol Lee, Du-seong Chang, and Jiwon Seo. 2024 · 2024
Closest in time.
SuffixDecoding: A Model-Free Approach to Speeding Up Large Language Model Inference
Gabriele Oliaro, Zhihao Jia, Daniel Campos, and Aurick Qiao. 2024 · 2024
Closest in time.
Marconi: Prefix caching for the era of hybrid llms
Rui Pan, Zhuang Wang, Zhen Jia, Can Karakus, Luca Zancato, Tri Dao, Yida Wang, and Ravi Netravali. 2024 · 2024
Closest in time.
Splitwise: Efficient generative llm inference using phase splitting. In Proc. of ISCA 2024 . 118–132
Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini. 2024 · 2024
Closest in time.
ConServe: Harvesting GPUs for Low-Latency and High-Throughput Large Language Model Serving
Yifan Qiao, Shu Anzai, Shan Yu, Haoran Ma, Yang Wang, Miryung Kim, and Harry Xu. 2024 · 2024
Closest in time.
Mooncake: A kvcache-centric disaggregated architecture for llm serving
Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, Yongwei Wu, Weimin Zheng, and Xinran Xu. 2024 · 2024
Closest in time.
Hybridflow: A flexible and efficient rlhf framework
Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, and Chuan Wu. 2024b · 2024
Closest in time.
Fairness in serving large language models. In Proc. of OSDI 2024 . 965–988
Ying Sheng, Shiyi Cao, Dacheng Li, Banghua Zhu, Zhuohan Li, Danyang Zhuo, Joseph E Gonzalez, and Ion Stoica. 2024a · 2024
Closest in time.
Llumnix: Dynamic scheduling for large language model serving. In Proc. of OSDI 2024 . 173–191
Biao Sun, Ziming Huang, Hanyu Zhao, Wencong Xiao, Xinyi Zhang, Yong Li, and Wei Lin. 2024b · 2024
Closest in time.
TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding. In COLM 2024
Hanshi Sun, Zhuoming Chen, Xinyu Yang, Yuandong Tian, and Beidi Chen. 2024a · 2024
Closest in time.
BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems
Yuxin Wang, Yuhan Chen, Zeyu Li, Xueze Kang, Zhenheng Tang, Xin He, Rui Guo, Xin Wang, Qiang Wang, Amelie Chi Zhou, et al · 2024
Closest in time.
Loongserve: Efficiently serving long-context large language models with elastic sequence parallelism. In Proc. of SOSP 2024 . 640–654
Bingyang Wu, Shengyu Liu, Yinmin Zhong, Peng Sun, Xuanzhe Liu, and Xin Jin. 2024b · 2024
Closest in time.
dLoRA: Dynamically orchestrating requests and adapters for LoRA LLM serving. In Proc. of OSDI 2024 . 911–927
Bingyang Wu, Ruidong Zhu, Zili Zhang, Peng Sun, Xuanzhe Liu, and Xin Jin. 2024c · 2024
Closest in time.
Mirage: A Multi-Level Superoptimizer for Tensor Programs
Mengdi Wu, Xinhao Cheng, Shengyu Liu, Chunan Shi, Jianan Ji, Kit Ao, Praveen Velliengiri, Xupeng Miao, Oded Padon, and Zhihao Jia. 2024a · 2024
Closest in time.
AI Metropolis: Scaling Large Language Model-based Multi-Agent Simulation with Out-of-order Execution
Zhiqiang Xie, Hao Kang, Ying Sheng, Tushar Krishna, Kayvon Fatahalian, and Christos Kozyrakis. 2024 · 2024
Closest in time.
Context Parallelism for Scalable Million-Token Inference
Amy Yang, Jingyi Yang, Aya Ibrahim, Xinfeng Xie, Bangsheng Tang, Grigory Sizov, Jeremy Reizenstein, Jongsoo Park, and Jianyu Huang. 2024 · 2024
Closest in time.
Large language model cascades with mixture of thoughts representations for cost-efficient reasoning. In Proc. of ICLR 2024
Murong Yue, Jie Zhao, Min Zhang, Du Liang, and Ziyu Yao. 2024 · 2024
Closest in time.
Pqcache: Product quantization-based kvcache for long context llm inference
Hailin Zhang, Xiaodong Ji, Yilin Chen, Fangcheng Fu, Xupeng Miao, Xiaonan Nie, Weipeng Chen, and Bin Cui. 2024a · 2024
Closest in time.
Draft& Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding. In Proc. of ACL 2024 . 11263–11282
Jun Zhang, Jue Wang, Huan Li, Lidan Shou, Ke Chen, Gang Chen, and Sharad Mehrotra. 2024b · 2024
Closest in time.
Accelerating iterative retrieval-augmented language model serving with speculation. In Proc. of ICML 2024
Zhihao Zhang, Alan Zhu, Lijie Yang, Yihua Xu, Lanting Li, Phitchaya Mangpo Phothilimthana, and Zhihao Jia. 2024c · 2024
Closest in time.
Ouroboros: Generating Longer Drafts Phrase by Phrase for Faster Speculative Decoding. In Proc. of EMNLP 2024 . 13378–13393
Weilin Zhao, Yuxiang Huang, Xu Han, Wang Xu, Chaojun Xiao, Xinrong Zhang, Yewei Fang, Kaihuo Zhang, Zhiyuan Liu, and Maosong Sun. 2024a · 2024
Closest in time.
Atom: Low-bit quantization for efficient and accurate llm serving
Yilong Zhao, Chien-Yu Lin, Kan Zhu, Zihao Ye, Lequn Chen, Size Zheng, Luis Ceze, Arvind Krishnamurthy, Tianqi Chen, and Baris Kasikci. 2024b · 2024
Closest in time.
DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving. In Proc. of OSDI 2024 . 193–210
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang. 2024a · 2024
Closest in time.
Rlhfuse: Efficient rlhf training for large language models with inter-and intra-stage fusion
Yinmin Zhong, Zili Zhang, Bingyang Wu, Shengyu Liu, Yukun Chen, Changyi Wan, Hanpeng Hu, Lei Xia, Ranchen Ming, Yibo Zhu, et al · 2024
Closest in time.
Fast state restoration in LLM serving with hcache. In EuroSys 2025
Shiwei Gao, Youmin Chen, and Jiwu Shu. 2025 · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
SpecServe: Efficient and SLO-Aware Large Language Model Serving with Adaptive Speculative Decoding
Kaiyu Huang, Hao Wu, Zhubo Shi, Han Zou, Minchen Yu, and Qingjiang Shi. 2025 · 2025
Closest in time.
GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline Parallelism. In Proc. of ASPLOS 2025 . 557–571
Byungsoo Jeon, Mengdi Wu, Shiyi Cao, Sunghyun Kim, Sunghyun Park, Neeraj Aggarwal, Colin Unger, Daiyaan Arfeen, Peiyuan Liao, Xupeng Miao, et al · 2025
Closest in time.
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
Youhe Jiang, Fangcheng Fu, Xiaozhe Yao, Guoliang He, Xupeng Miao, Ana Klimovic, Bin Cui, Binhang Yuan, and Eiko Yoneki. 2025 · 2025
Closest in time.
Relax: Composable Abstractions for End-to-End Dynamic Machine Learning. In Proc. of ASPLOS 2025
Ruihang Lai, Junru Shao, Siyuan Feng, Steven S Lyubomirsky, Bohan Hou, Wuwei Lin, Zihao Ye, Hongyi Jin, et al · 2025
Closest in time.
AdaServe: SLO-Customized LLM Serving with Fine-Grained Speculative Decoding
Zikun Li, Zhuofu Chen, Remi Delacourt, Gabriele Oliaro, Zeyu Wang, Qinghan Chen, Shuhuai Lin, April Yang, Zhihao Zhang, Zhuoming Chen, et al · 2025
Closest in time.
Jingyu Liu, Beidi Chen, and Ce Zhang. 2025 · 2025
Closest in time.
Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow. In Proc. of ASPLOS 2025 . 586–602
Yixuan Mei, Yonghao Zhuang, Xupeng Miao, Juncheng Yang, Zhihao Jia, and Rashmi Vinayak. 2025 · 2025
Closest in time.
vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention. In Proc. of ASPLOS 2025 . 1133–1150
Ramya Prabhu, Ajay Nayak, Jayashree Mohan, Ramachandran Ramjee, and Ashish Panwar. 2025 · 2025
Closest in time.
Opt-tree: Speculative decoding with adaptive draft tree structure
Jikai Wang, Yi Su, Juntao Li, Qingrong Xia, Zi Ye, Xinyu Duan, Zhefeng Wang, and Min Zhang. 2025 · 2025
Closest in time.
LongSpec: Long-Context Speculative Decoding with Efficient Drafting and Verification
Penghui Yang, Cunxiao Du, Fengzhuo Zhang, Haonan Wang, Tianyu Pang, Chao Du, and Bo An. 2025 · 2025
Closest in time.
Flashinfer: Efficient and customizable attention engine for llm inference serving
Zihao Ye, Lequn Chen, Ruihang Lai, Wuwei Lin, Yineng Zhang, Stephanie Wang, Tianqi Chen, Baris Kasikci, Vinod Grover, Arvind Krishnamurthy, et al · 2025
Closest in time.
DeepEP: an efficient expert-parallel communication library
Chenggang Zhao, Shangyan Zhou, Liyue Zhang, Chengqi Deng, Zhean Xu, Yuxuan Liu, Kuai Yu, Jiashi Li, and Liang Zhao. 2025 · 2025
Closest in time.