Fetching the paper…
Reading the bibliography…
Nowadays, large language models (LLMs) are published as a service and can be accessed by various applications via APIs, also known as language-model-as-a-service (LMaaS).
F. Pedregosa, G. Varoquaux, A. Gramfort et al. , “Scikit-learn: Machine learning in python,” the Journal of machine Learning research , vol. 12, pp. 2825–2830, 2011
2011
Earlier work this paper cites.
M. Freitag and Y. Al-Onaizan, “Beam search strategies for neural machine translation,” in Proceedings of the First Workshop on Neural Machine Translation . Association for Computational Linguistics, 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar et al. , “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
C. L. Goues, M. Pradel, and A. Roychoudhury, “Automated program repair,” Communications of the ACM , vol. 62, no. 12, pp. 56–65, 2019
2019
Earlier work this paper cites.
N. Reimers and I. Gurevych, “Sentence-bert: Sentence embeddings using siamese bert-networks,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) . Association for Computational Linguistics, 2019
2019
Earlier work this paper cites.
B. Yang, Z. Liping, and Z. Fengrong, “A survey on research of code comment,” in Proceedings of the 2019 3rd International Conference on Management Engineering, Software Engineering and Service Sciences , 2019, pp. 45–51
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
Z. Wang, J. Wohlwend, and T. Lei, “Structured pruning of large language models,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2020, pp. 6151–6162
2020
Earlier work this paper cites.
T. Wolf, L. Debut, V. Sanh et al. , “Transformers: State-of-the-art natural language processing,” in Proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations , 2020, pp. 38–45
2020
Earlier work this paper cites.
F. Stahlberg, “Neural machine translation: A review,” Journal of Artificial Intelligence Research , vol. 69, pp. 343–418, 2020
2020
Earlier work this paper cites.
D. Xu, I. E.-H. Yen, J. Zhao, and Z. Xiao, “Rethinking network pruning–under the pre-train and fine-tune paradigm,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , 2021, pp. 2376–2382
2021
Earlier work this paper cites.
F. Stahlberg and S. Kumar, “Synthetic data generation for grammatical error correction with tagged corruption models,” in Proceedings of the 16th Workshop on Innovative Use of NLP for Building Educational Applications , 2021, pp. 37–47
2021
Earlier work this paper cites.
S. Lu, D. Guo, S. Ren et al. , “Codexglue: A machine learning benchmark dataset for code understanding and generation,” in Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track , 2021
2021
Earlier work this paper cites.
M. Yasunaga and P. Liang, “Break-it-fix-it: Unsupervised learning for program repair,” in International Conference on Machine Learning . PMLR, 2021, pp. 11 941–11 952
2021
Earlier work this paper cites.
Y. Wang, Y. Wang, K. Dang et al. , “A comprehensive survey of grammatical error correction,” ACM Transactions on Intelligent Systems and Technology (TIST) , vol. 12, no. 5, pp. 1–51, 2021
2021
Earlier work this paper cites.
D. Dale, A. Voronov, D. Dementieva et al. , “Text detoxification using large pre-trained neural models,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing , 2021, pp. 7979–7996
2021
Earlier work this paper cites.
J. Fang, Y. Yu, C. Zhao, and J. Zhou, “Turbotransformers: an efficient gpu serving system for transformer models,” in Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming , 2021, pp. 389–402
2021
Earlier work this paper cites.
S. Zhang, S. Roller, N. Goyal et al. , “Opt: Open pre-trained transformer language models,” 2022
2022
Earlier work this paper cites.
Z. Du, Y. Qian, X. Liu et al. , “Glm: General language model pretraining with autoregressive blank infilling,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2022, pp. 320–335
2022
Earlier work this paper cites.
T. Sun, Y. Shao, H. Qian et al. , “Black-box tuning for language-model-as-a-service,” in International Conference on Machine Learning . PMLR, 2022, pp. 20 841–20 855
2022
Cited alongside, same era.
G.-I. Yu, J. S. Jeong, G.-W. Kim et al. , “Orca: A distributed serving system for { \{ Transformer-Based } \} generative models,” in 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22) , 2022, pp. 521–538
2022
Cited alongside, same era.
Z. Yao, R. Yazdani Aminabadi, M. Zhang et al. , “Zeroquant: Efficient and affordable post-training quantization for large-scale transformers,” Advances in Neural Information Processing Systems , vol. 35, pp. 27 168–27 183, 2022
2022
Cited alongside, same era.
R. Xu, F. Luo, C. Wang et al. , “From dense to sparse: Contrastive pruning for better pre-trained language model compression,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 10, 2022, pp. 11 547–11 555
2022
Cited alongside, same era.
Z. Liu, J. Wang, T. Dao et al. , “Deja vu: Contextual sparsity for efficient llms at inference time,” in International Conference on Machine Learning . PMLR, 2023, pp. 22 137–22 176
2023
Later among the works it cites.
M. Sun, Z. Liu, A. Bair, and J. Z. Kolter, “A simple and effective pruning approach for large language models,” in The Twelfth International Conference on Learning Representations , 2023
2023
Later among the works it cites.
“Wmt18,” https://huggingface.co/datasets/wmt18
2023
Later among the works it cites.
F. Bang, “Gptcache: An open-source semantic cache for llm applications enabling faster answers and cost savings,” in Proceedings of the 3rd Workshop for Natural Language Processing Open Source Software (NLP-OSS 2023) , 2023, pp. 212–218
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
V. Logacheva, D. Dementieva, S. Ustyantsev et al. , “Paradetox: Detoxification with parallel data,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2022, pp. 6804–6818
2022
Cited alongside, same era.
R. Y. Aminabadi, S. Rajbhandari, A. A. Awan et al. , “Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,” in Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis , 2022, pp. 1–15
2022
Cited alongside, same era.
F. Feng, Y. Yang, D. Cer et al. , “Language-agnostic bert sentence embedding,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2022, pp. 878–891
2022
Cited alongside, same era.
T. Dao, D. Fu, S. Ermon et al. , “Flashattention: Fast and memory-efficient exact attention with io-awareness,” Advances in Neural Information Processing Systems , vol. 35, pp. 16 344–16 359, 2022
2022
Cited alongside, same era.
M. Zhu, K. Suresh, and C. K. Reddy, “Multilingual code snippets training for program translation,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 10, 2022, pp. 11 783–11 790
2022
Cited alongside, same era.
B. Fu, F. Chen, P. Li, and D. Zeng, “Tcb: Accelerating transformer inference services with request concatenation,” in Proceedings of the 51st International Conference on Parallel Processing , 2022, pp. 1–11
2022
Cited alongside, same era.
H. Touvron, T. Lavril, G. Izacard et al. , “Llama: Open and efficient foundation language models,” 2023
2023
Cited alongside, same era.
H. Touvron, L. Martin, K. Stone et al. , “Llama 2: Open foundation and fine-tuned chat models,” 2023
2023
Cited alongside, same era.
2023
Later among the works it cites.
T. Dao, “Flashattention-2: Faster attention with better parallelism and work partitioning,” in The Twelfth International Conference on Learning Representations , 2023
2023
Later among the works it cites.
Y. Leviathan, M. Kalman, and Y. Matias, “Fast inference from transformers via speculative decoding,” in International Conference on Machine Learning . PMLR, 2023, pp. 19 274–19 286
2023
Later among the works it cites.
H. Xia, T. Ge, P. Wang et al. , “Speculative decoding: Exploiting speculative execution for accelerating seq2seq generation,” in Findings of the Association for Computational Linguistics: EMNLP 2023 , 2023, pp. 3909–3925
2023
Later among the works it cites.
“Baichuan2-7b-chat,” https://huggingface.co/baichuan-inc/Baichuan2-7B-Chat
2024
Closest in time.
“Deepspeed-fastgen,” https://github.com/microsoft/DeepSpeed-MII
2024
Closest in time.
J. Chee, Y. Cai, V. Kuleshov et al. , “Quip: 2-bit quantization of large language models with guarantees,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
S. Ge, Y. Zhang, L. Liu et al. , “Model tells you what to discard: Adaptive kv cache compression for llms,” in The Twelfth International Conference on Learning Representations , 2024
2024
Closest in time.
Z. Zhang, Y. Sheng, T. Zhou et al. , “H2o: Heavy-hitter oracle for efficient generative inference of large language models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
Z. Liu, A. Desai, F. Liao et al. , “Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
Z. Sun, A. T. Suresh, J. H. Ro et al. , “Spectr: Fast speculative decoding via optimal transport,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
S. Kim, K. Mangalam, S. Moon et al. , “Speculative decoding with big little decoder,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
Z. Zheng, X. Ren, F. Xue et al. , “Response length perception and sequence scheduling: An llm-empowered llm inference pipeline,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
Y. Jin, C.-F. Wu, D. Brooks, and G.-Y. Wei, “ s 3 s^{3} : Increasing gpu utilization during generative inference for higher throughput,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.