Fetching the paper…
Reading the bibliography…
The rapid advancement of large language models (LLMs) has significantly improved code completion tasks, yet the trade-off between accuracy and computational cost remains a critical challenge.
Monotone piecewise cubic interpolation
Frederick N Fritsch and Ralph E Carlson · 1980
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, et al · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, et al · 2021
Earlier work this paper cites.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, et al · 2021
Earlier work this paper cites.
Measuring coding challenge competence with apps
Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, et al · 2021
Earlier work this paper cites.
Accelerate: Training and inference at scale made simple, efficient and adaptable
Sylvain Gugger, Lysandre Debut, Thomas Wolf, Philipp Schmid, Zachary Mueller, et al · 2022
Earlier work this paper cites.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, et al · 2023
Earlier work this paper cites.
Wizardcoder: Empowering code large language models with evol-instruct
Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, et al · 2023
Earlier work this paper cites.
Fast inference from transformers via speculative decoding
Yaniv Leviathan, Matan Kalman, and Yossi Matias · 2023
Cited alongside, same era.
The program testing ability of large language models for code
Weimin Xiong, Yiwen Guo, and Hao Chen · 2023
Cited alongside, same era.
Gpu cloud pricing
RunPod · 2023
Cited alongside, same era.
From words to watts: Benchmarking the energy costs of large language model inference
Siddharth Samsi, Dan Zhao, Joseph McDonald, Baolin Li, Adam Michaleas, et al · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, et al · 2023
Cited alongside, same era.
Self-improvement in language models: The sharpening mechanism
Audrey Huang, Adam Block, Dylan J. Foster, Dhruv Rohatgi, Cyril Zhang, et al · 2024
Closest in time.
A survey on model compression for large language models
Xunyu Zhu, Jian Li, Yong Liu, Can Ma, and Weiping Wang · 2024
Closest in time.
Duquant: Distributing outliers via dual transformation makes stronger quantized LLMs
Haokun Lin, Haobo Xu, Yichen Wu, Jingzhi Cui, Yingtao Zhang, et al · 2024
Closest in time.
Threshold-driven pruning with segmented maximum term weights for approximate cluster-based sparse retrieval
Yifan Qiao, Parker Carlson, Shanxiu He, Yingrui Yang, and Tao Yang · 2024
Closest in time.
AMR-evol: Adaptive modular response evolution elicits better knowledge distillation for large language models in code generation
Ziyang Luo, Xin Li, Hongzhan Lin, Jing Ma, and Lidong Bing · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sahil Chaudhary · 2023
Cited alongside, same era.
Deepseek-coder: When the large language model meets programming – the rise of code intelligence
Daya Guo, Qihao Zhu, Dejian Yang, and Zhenda Xie · 2024
Cited alongside, same era.
Code llama: Open foundation models for code
Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, et al · 2024
Cited alongside, same era.
Codejudge: Evaluating code generation with large language models
Weixi Tong and Tianyi Zhang · 2024
Cited alongside, same era.
Codet: Code generation with generated tests
Bei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan, Zeqi Lin, et al
Cited in the paper.
Frugalgpt: How to use large language models while reducing cost and improving performance
Lingjiao Chen, Matei Zaharia, and James Zou
Cited in the paper.
Scalable efficient training of large language models with low-dimensional projected attention
Xingtai Lv, Ning Ding, Kaiyan Zhang, Ermo Hua, Ganqu Cui, et al · 2024
Closest in time.
Large language model cascades with mixture of thought representations for cost-efficient reasoning
Murong Yue, Jie Zhao, Min Zhang, Liang Du, and Ziyu Yao · 2024
Closest in time.
Qwen2.5-coder technical report
Binyuan Hui, Jian Yang, Zeyu Cui, Jiaxi Yang, Dayiheng Liu, et al · 2024
Closest in time.
Bass: Batched attention-optimized speculative sampling
Haifeng Qian, Sujan Kumar Gonugondla, Sungsoo Ha, Mingyue Shang, Sanjay Krishna Gouda, et al · 2024
Closest in time.