Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have significantly advanced software engineering (SE) tasks, with prompt engineering techniques enhancing their performance in code-related areas.
Individual comparisons by ranking methods
Frank Wilcoxon. 1945 · 1945
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics . 311–318
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
The probabilistic relevance framework: BM25 and beyond
Stephen Robertson, Hugo Zaragoza, et al · 2009
Earlier work this paper cites.
CodeSearchNet challenge: Evaluating the state of semantic code search
Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt. 2019 · 2019
Earlier work this paper cites.
Codebleu: a method for automatic evaluation of code synthesis
Shuo Ren, Daya Guo, Shuai Lu, Long Zhou, Shujie Liu, Duyu Tang, Neel Sundaresan, Ming Zhou, Ambrosio Blanco, and Shuai Ma. 2020 · 2020
Earlier work this paper cites.
Unsupervised translation of programming languages
Baptiste Roziere, Marie-Anne Lachaux, Lowik Chanussot, and Guillaume Lample. 2020 · 2020
Earlier work this paper cites.
Evaluating Large Language Models Trained on Code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. 2021 · 2021
Earlier work this paper cites.
CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation
Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin B. Clement, Dawn Drain, Daxin Jiang, Duyu Tang, Ge Li, Lidong Zhou, Linjun Shou, Long Zhou, Michele Tufano, Ming Gong, Ming Zhou, Nan Duan, Neel Sundaresan, Shao Kun Deng, Shengyu Fu, and Shujie Liu. 2021 · 2021
Earlier work this paper cites.
ChatGPT: A Large-Scale Chatbot Model
OpenAI. 2021 · 2021
Earlier work this paper cites.
Carbon Emissions and Large Neural Network Training
David Patterson, Joseph Gonzalez, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean. 2021 · 2021
Earlier work this paper cites.
On the evaluation of commit message generation models: an experimental study. In 2021 IEEE International Conference on Software Maintenance and Evolution (ICSME) . IEEE, 126–136
Wei Tao, Yanlin Wang, Ensheng Shi, Lun Du, Shi Han, Hongyu Zhang, Dongmei Zhang, and Wenqiang Zhang. 2021 · 2021
Earlier work this paper cites.
eco2AI: Carbon Emissions Tracking of Machine Learning Models as the First Step Towards Sustainable AI
Nikita N Lazarev, Andrey N Zakharenko, Oleg A Korovin, Daria V Plosskaya, Vladimir S Dimitrov, Ivan V Akhripkin, Ivan V Pavlov, Ivan V Oseledets, Ivan S Barsola, Alexander A Egorov, et al · 2022
Earlier work this paper cites.
RACE: Retrieval-augmented Commit Message Generation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, Abu Dhabi, United Arab Emirates, 5520–5530
Ensheng Shi, Yanlin Wang, Wei Tao, Lun Du, Hongyu Zhang, Shi Han, Dongmei Zhang, and Hongbin Sun. 2022 · 2022
Earlier work this paper cites.
What makes a good commit message?. In Proceedings of the 44th International Conference on Software Engineering . 2389–2401
Yingchen Tian, Yuxia Zhang, Klaas-Jan Stol, Lin Jiang, and Hui Liu. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022 · 2022
Earlier work this paper cites.
Summarize and Generate to Back-translate: Unsupervised Translation of Programming Languages. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics . 1528–1542
Wasi Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2023 · 2023
Earlier work this paper cites.
LLMCarbon: Modeling the End-to-End Carbon Footprint of Large Language Models
Anonymous. 2023 · 2023
Earlier work this paper cites.
From Commit Message Generation to History-Aware Commit Message Completion. In 2023 38th IEEE/ACM International Conference on Automated Software Engineering (ASE) . IEEE, 723–735
Aleksandra Eliseeva, Yaroslav Sokolov, Egor Bogomolov, Yaroslav Golubev, Danny Dig, and Timofey Bryksin. 2023 · 2023
Earlier work this paper cites.
What makes good in-context demonstrations for code intelligence tasks with llms?. In 2023 38th IEEE/ACM International Conference on Automated Software Engineering (ASE) . IEEE, 761–773
Shuzheng Gao, Xin-Cheng Wen, Cuiyun Gao, Wenxuan Wang, Hongyu Zhang, and Michael R Lyu. 2023 · 2023
Earlier work this paper cites.
Agentcoder: Multi-agent-based code generation with iterative testing and optimisation
Dong Huang, Jie M Zhang, Michael Luck, Qingwen Bu, Yuhao Qing, and Heming Cui. 2023 · 2023
Cited alongside, same era.
Timeshifting Strategies for Carbon-Efficient Long-Running Large Language Model Training
Akshaya Jagannadharao, Nicole Beckage, Dawn Nafus, and Scott Chamberlin. 2023 · 2023
Cited alongside, same era.
Explainable automated debugging via large language model-driven scientific debugging
Sungmin Kang, Bei Chen, Shin Yoo, and Jian-Guang Lou. 2023 · 2023
Cited alongside, same era.
Benchmarking cognitive biases in large language models as evaluators
Ryan Koo, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim, and Dongyeop Kang. 2023 · 2023
Cited alongside, same era.
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
DeepSeek-AI. 2024 · 2024
Closest in time.
Self-collaboration code generation via chatgpt
Yihong Dong, Xue Jiang, Zhi Jin, and Ge Li. 2024 · 2024
Closest in time.
A quantitative and qualitative evaluation of LLM-based explainable fault localization
Sungmin Kang, Gabin An, and Shin Yoo. 2024 · 2024
Closest in time.
Language models can solve computer tasks
Geunwoo Kim, Pierre Baldi, and Stephen McAleer. 2024 · 2024
Closest in time.
Large language model-based agents for software engineering: A survey
Junwei Liu, Kaixin Wang, Yixuan Chen, Xin Peng, Zhenpeng Chen, Lingming Zhang, and Yiling Lou. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bo Lin, Shangwen Wang, Zhongxin Liu, Yepang Liu, Xin Xia, and Xiaoguang Mao. 2023 · 2023
Cited alongside, same era.
Large language models are effective text rankers with pairwise ranking prompting
Zhen Qin, Rolf Jagerman, Kai Hui, Honglei Zhuang, Junru Wu, Le Yan, Jiaming Shen, Tianqi Liu, Jialu Liu, Donald Metzler, et al · 2023
Cited alongside, same era.
Sentiment analysis through llm negotiations
Xiaofei Sun, Xiaoya Li, Shengyu Zhang, Shuhe Wang, Fei Wu, Jiwei Li, Tianwei Zhang, and Guoyin Wang. 2023 · 2023
Cited alongside, same era.
Element-aware Summarization with Large Language Models: Expert-aligned Evaluation and Chain-of-Thought Method. In The 61st Annual Meeting Of The Association For Computational Linguistics
Yiming Wang, Zhuosheng Zhang, and Rui Wang. 2023 · 2023
Cited alongside, same era.
Keep the Conversation Going: Fixing 162 out of 337 bugs for $0.42 each using ChatGPT
Chunqiu Steven Xia and Lingming Zhang. 2023 · 2023
Cited alongside, same era.
Expertprompting: Instructing large language models to be distinguished experts
Benfeng Xu, An Yang, Junyang Lin, Quan Wang, Chang Zhou, Yongdong Zhang, and Zhendong Mao. 2023b · 2023
Cited alongside, same era.
Energy Efficiency of Training Neural Network Architectures: An Empirical Study
Yinlena Xu, Silverio Martínez-Fernández, Matias Martinez, and Xavier Franch. 2023a · 2023
Cited alongside, same era.
OpenAI API Pricing
2024 · 2024
Cited alongside, same era.
Wei Ma, Daoyuan Wu, Yuqiang Sun, Tianwen Wang, Shangqing Liu, Jian Zhang, Yue Xue, and Yang Liu. 2024 · 2024
Closest in time.
Large language model (llm) ai text generation detection based on transformer deep learning algorithm
Yuhong Mo, Hao Qin, Yushan Dong, Ziyi Zhu, and Zhenglin Li. 2024 · 2024
Closest in time.
GPT-4o: Optimized GPT-4 Language Model
OpenAI. 2024b · 2024
Closest in time.
Learning to Reason with LLMs
OpenAI. 2024c · 2024
Closest in time.
OpenAI o1-mini
OpenAI. 2024d · 2024
Closest in time.
Lost in translation: A study of bugs introduced by large language models while translating code. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering . 1–13
Rangeet Pan, Ali Reza Ibrahimzada, Rahul Krishna, Divya Sankar, Lambert Pouguem Wassi, Michele Merler, Boris Sobolev, Raju Pavuluri, Saurabh Sinha, and Reyhaneh Jabbarvand. 2024 · 2024
Closest in time.
Can llms master math? investigating large language models on math stack exchange. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval . 2316–2320
Ankit Satpute, Noah Gießing, André Greiner-Petter, Moritz Schubotz, Olaf Teschke, Akiko Aizawa, and Bela Gipp. 2024 · 2024
Closest in time.
A survey on large language models for recommendation
Likang Wu, Zhi Zheng, Zhaopeng Qiu, Hao Wang, Hongchao Gu, Tingjia Shen, Chuan Qin, Chen Zhu, Hengshu Zhu, Qi Liu, et al · 2024
Closest in time.
Evaluating the Effectiveness of LLM-Evaluators (aka LLM-as-Judge)
Ziyou Yan. 2024 · 2024
Closest in time.
Exploring and unleashing the power of large language models in automated code translation
Zhen Yang, Fang Liu, Zhongxing Yu, Jacky Wai Keung, Jia Li, Shuo Liu, Yifan Hong, Xiaoxue Ma, Zhi Jin, and Ge Li. 2024 · 2024
Closest in time.
A Survey on Recent Advances in LLM-Based Multi-turn Dialogue Systems
Zihao Yi, Jiarui Ouyang, Yuwen Liu, Tianhao Liao, Zhe Xu, and Ying Shen. 2024 · 2024
Closest in time.
Fight Fire with Fire: How Much Can We Trust ChatGPT on Source Code-Related Tasks?
Xiao Yu, Lei Liu, Xing Hu, Jacky Wai Keung, Jin Liu, and Xin Xia. 2024 · 2024
Closest in time.
Debug like a Human: A Large Language Model Debugger via Verifying Runtime Execution Step by Step. In Findings of the Association for Computational Linguistics ACL 2024 . 851–870
Li Zhong, Zilong Wang, and Jingbo Shang. 2024 · 2024
Closest in time.
Source Code Summarization in the Era of Large Language Models. In 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE) . IEEE Computer Society, 419–431
Weisong Sun, Yun Miao, Yuekang Li, Hongyu Zhang, Chunrong Fang, Yi Liu, Gelei Deng, Yang Liu, and Zhenyu Chen. 2024 · 2025
Closest in time.