Fetching the paper…
Reading the bibliography…
The "pretrain-then-finetune" paradigm is commonly adopted in the deployment of large language models.
An admission control algorithm for predictive real-time service
Jamin, S., Shenker, S., Zhang, L., and Clark, D. D · 1993
Earlier work this paper cites.
A statistical admission control algorithm for multimedia servers
Vin, H., Goyal, P., and Goyal, A · 1994
Earlier work this paper cites.
Distributed call admission control in mobile/wireless networks
Naghshineh, M. and Schwartz, M · 1996
Earlier work this paper cites.
Clipper: A low-latency online prediction serving system
Crankshaw, D., Wang, X., Zhou, G., Franklin, M. J., Gonzalez, J. E., and Stoica, I · 2017
Earlier work this paper cites.
Tensorflow-serving: Flexible, high-performance ml serving
Olston, C., Fiedel, N., Gorovoy, K., Harmsen, J., Lao, L., Li, F., Rajashekhar, V., Ramesh, S., and Soyke, J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Blockwise parallel decoding for deep autoregressive models
Stern, M., Shazeer, N., and Uszkoreit, J · 2018
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S · 2019
Earlier work this paper cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, D., Chen, M., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., et al · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Kenton, J. D. M.-W. C. and Toutanova, L. K · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Earlier work this paper cites.
Fast transformer decoding: One write-head is all you need
Shazeer, N · 2019
Earlier work this paper cites.
Nexus: A gpu cluster engine for accelerating dnn-based video analysis
Shen, H., Chen, L., Jin, Y., Zhao, L., Kong, B., Philipose, M., Krishnamurthy, A., and Sundaram, R · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 2019
Earlier work this paper cites.
Triton: an intermediate language and compiler for tiled neural network computations
Tillet, P., Kung, H.-T., and Cox, D · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Inferline: latency-aware provisioning and scaling for prediction serving pipelines
Crankshaw, D., Sela, G.-E., Mo, X., Zumar, C., Stoica, I., Gonzalez, J., and Tumanov, A · 2020
Earlier work this paper cites.
Serving { \{ DNNs } \} like clockwork: Performance predictability from the bottom up
Gujarati, A., Karimi, R., Alzayat, S., Hao, W., Kaufmann, A., Vigfusson, Y., and Mace, J · 2020
Earlier work this paper cites.
Turbotransformers: an efficient gpu serving system for transformer models
Fang, J., Yu, Y., Zhao, C., and Zhou, J · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X. L. and Liang, P · 2021
Cited alongside, same era.
P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks
Liu, X., Ji, K., Fu, Y., Tam, W. L., Du, Z., Yang, Z., and Tang, J · 2021
Cited alongside, same era.
Efficient large-scale language model training on gpu clusters using megatron-lm
Narayanan, D., Shoeybi, M., Casper, J., LeGresley, P., Patwary, M., Korthikanti, V., Vainbrand, D., Kashinkunti, P., Bernauer, J., Catanzaro, B., et al · 2021
Cited alongside, same era.
Lightseq: A high performance inference library for transformers
Wang, X., Xiong, Y., Wei, Y., Wang, M., and Li, L · 2021
Cited alongside, same era.
Adaptive budget allocation for parameter-efficient fine-tuning
Zhang, Q., Chen, M., Bukharin, A., He, P., Cheng, Y., Chen, W., and Zhao, T · 2022
Later among the works it cites.
Alpa: Automating inter-and intra-operator parallelism for distributed deep learning
Zheng, L., Li, Z., Zhang, H., Zhuang, Y., Chen, Z., Huang, Y., Wang, Y., Xu, Y., Zhuo, D., Xing, E. P., et al · 2022
Later among the works it cites.
{ \{ PetS } \} : A unified framework for { \{ Parameter-Efficient } \} transformers serving
Zhou, Z., Wei, X., Zhang, J., and Sun, G · 2022
Later among the works it cites.
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al · 2023
Closest in time.
Potentials of multitenancy fine-tuned llm serving
Chen, L · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Cited alongside, same era.
Deepspeed- inference: Enabling efficient inference of transformer models at unprecedented scale
Aminabadi, R. Y., Rajbhandari, S., Awan, A. A., Li, C., Li, D., Zheng, E., Ruwase, O., Smith, S., Zhang, M., Rasley, J., and He, Y · 2022
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Cited alongside, same era.
Dvabatch: Diversity-aware multi-entry multi-exit batching for efficient processing of dnn services on gpus
Cui, W., Zhao, H., Chen, Q., Wei, H., Li, Z., Zeng, D., Li, C., and Guo, M · 2022
Cited alongside, same era.
Llm.int8(): 8-bit matrix multiplication for transformers at scale
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L · 2022
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2022
Cited alongside, same era.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2022
Cited alongside, same era.
Chen, L., Ye, Z., Wu, Y., Zhuo, D., Ceze, L., and Krishnamurthy, A · 2023
Closest in time.
Flashattention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2023
Closest in time.
Massive language models can be accurately pruned in one-shot
Frantar, E. and Alistarh, D · 2023
Closest in time.
Reducing activation recomputation in large transformer models
Korthikanti, V. A., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B · 2023
Closest in time.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Closest in time.
{ \{ AlpaServe } \} : Statistical multiplexing with model parallelism for deep learning serving
Li, Z., Zheng, L., Zhong, Y., Liu, V., Sheng, Y., Jin, X., Huang, Y., Chen, Z., Zhang, H., Gonzalez, J. E., et al · 2023
Closest in time.
Awq: Activation-aware weight quantization for llm compression and acceleration
Lin, J., Tang, J., Tang, H., Yang, S., Dang, X., and Han, S · 2023
Closest in time.
Gpt understands, too
Liu, X., Zheng, Y., Du, Z., Ding, M., Qian, Y., Yang, Z., and Tang, J · 2023
Closest in time.
Miao, X., Oliaro, G., Zhang, Z., Cheng, X., Wang, Z., Wong, R. Y. Y., Chen, Z., Arfeen, D., Abhyankar, R., and Jia, Z · 2023
Closest in time.
Lightllm: Python-based llm inference and serving framework
ModelTC · 2023
Closest in time.
Fastertransformer
NVIDIA · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Flexgen: High-throughput generative inference of large language models with a single GPU
Sheng, Y., Zheng, L., Yuan, B., Li, Z., Ryabinin, M., Chen, B., Liang, P., Ré, C., Stoica, I., and Zhang, C · 2023
Closest in time.
Smoothquant: Accurate and efficient post-training quantization for large language models
Xiao, G., Lin, J., Seznec, M., Wu, H., Demouth, J., and Han, S · 2023
Closest in time.