Fetching the paper…
Reading the bibliography…
Recently, Large Language Models (LLMs) have been increasingly used to automate SE tasks such as code generation and summarization.
Cognitive task analysis
Jan Maarten Schraagen, Susan F Chipman, and Valerie L Shalin. 2000 · 2000
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics . 311–318
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries. In Text summarization branches out . 74–81
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization . 65–72
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Software quality engineering: testing, quality assurance, and quantifiable improvement
Jeff Tian. 2005 · 2005
Earlier work this paper cites.
Structured and semistructured interviews
Daniel L Segal, Frederick L Coolidge, Alisa O’Riley, and Benjamin A Heinz. 2006 · 2006
Earlier work this paper cites.
Benefits and barriers of user evaluation in software engineering research. In Proceedings of the 2011 ACM international conference on Object oriented programming systems languages and applications . 643–656
Raymond PL Buse, Caitlin Sadowski, and Westley Weimer. 2011 · 2011
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. 2015 · 2015
Earlier work this paper cites.
Deep learning . Vol. 1
Ian Goodfellow, Yoshua Bengio, Aaron Courville, and Yoshua Bengio. 2016 · 2016
Earlier work this paper cites.
chrF++: words helping character n-grams. In Proceedings of the second conference on machine translation . 612–618
Maja Popović. 2017 · 2017
Earlier work this paper cites.
Learning to mine aligned code and natural language pairs from stack overflow. In Proceedings of the 15th international conference on mining software repositories . 476–486
Pengcheng Yin, Bowen Deng, Edgar Chen, Bogdan Vasilescu, and Graham Neubig. 2018 · 2018
Earlier work this paper cites.
Think-aloud protocols
Lawrence Jun Zhang and Donglan Zhang. 2019 · 2019
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 2019
Earlier work this paper cites.
Codebleu: a method for automatic evaluation of code synthesis
Shuo Ren, Daya Guo, Shuai Lu, Long Zhou, Shujie Liu, Duyu Tang, Neel Sundaresan, Ming Zhou, Ambrosio Blanco, and Shuai Ma. 2020 · 2020
Earlier work this paper cites.
Adversarial examples for models of code
Noam Yefet, Uri Alon, and Eran Yahav. 2020 · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Measuring coding challenge competence with apps (2021)
Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, et al · 2021
Earlier work this paper cites.
Reassessing automatic evaluation metrics for code summarization tasks. In Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering . 1105–1116
Devjeet Roy, Sarah Fakhoury, and Venera Arnaoudova. 2021 · 2021
Earlier work this paper cites.
BiasHeal: On-the-Fly Black-Box Healing of Bias in Sentiment Analysis Systems. In 2021 IEEE International Conference on Software Maintenance and Evolution (ICSME) . 644–648
Zhou Yang, Harshit Jain, Jieke Shi, Muhammad Hilmi Asyrofi, and David Lo. 2021 · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Earlier work this paper cites.
Code generation using machine learning: A systematic review
Enrique Dehaerne, Bappaditya Dey, Sandip Halder, Stefan De Gendt, and Wannes Meert. 2022 · 2022
Earlier work this paper cites.
Curiosity-driven and victim-aware adversarial policies. In Proceedings of the 38th Annual Computer Security Applications Conference . 186–200
Chen Gong, Zhou Yang, Yunpeng Bai, Jieke Shi, Arunesh Sinha, Bowen Xu, David Lo, Xinwen Hou, and Guoliang Fan. 2022 · 2022
Earlier work this paper cites.
Correlating automated and human evaluation of code documentation generation quality
Xing Hu, Qiuyuan Chen, Haoye Wang, Xin Xia, David Lo, and Thomas Zimmermann. 2022 · 2022
Earlier work this paper cites.
On the Influence of Biases in Bug Localization: Evaluation and Benchmark. In 2022 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER) . 128–139
Ratnadira Widyasari, Stefanus Agus Haryono, Ferdian Thung, Jieke Shi, Constance Tan, Fiona Wee, Jack Phan, and David Lo. 2022 · 2022
Earlier work this paper cites.
Revisiting Neuron Coverage Metrics and Quality of Deep Neural Networks. In 2022 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER) . 408–419
Zhou Yang, Jieke Shi, Muhammad Hilmi Asyrofi, and David Lo. 2022a · 2022
Earlier work this paper cites.
Claude (Oct 8 version) [Large language model]
Anthropic. 2023 · 2023
Earlier work this paper cites.
Out of the bleu: how should we assess quality of the code generation models?
Mikhail Evtikhiev, Egor Bogomolov, Yaroslav Sokolov, and Timofey Bryksin. 2023 · 2023
Earlier work this paper cites.
Large Language Models for Software Engineering: Survey and Open Problems. In 2023 IEEE/ACM International Conference on Software Engineering: Future of Software Engineering (ICSE-FoSE) . 31–53
Angela Fan, Beliz Gokkaya, Mark Harman, Mitya Lyubarskiy, Shubho Sengupta, Shin Yoo, and Jie M. Zhang. 2023 · 2023
Earlier work this paper cites.
Gptscore: Evaluate as you desire
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2023 · 2023
Earlier work this paper cites.
Impact of Code Language Models on Automated Program Repair. In Proceedings of the 45th International Conference on Software Engineering (Melbourne, Victoria, Australia) (ICSE ’23) . IEEE Press, 1430–1442
Nan Jiang, Kevin Liu, Thibaud Lutellier, and Lin Tan. 2023 · 2023
Cited alongside, same era.
InferFix: End-to-End Program Repair with LLMs. In Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (San Francisco, CA, USA) (ESEC/FSE 2023) . Association for Computing Machinery, New York, NY, USA, 1646–1656
Matthew Jin, Syed Shahriar, Michele Tufano, Xin Shi, Shuai Lu, Neel Sundaresan, and Alexey Svyatkovskiy. 2023 · 2023
Cited alongside, same era.
Neuro-symbolic models of human moral judgment: LLMs as automatic feature extractors
Joe Kwon, Sydney Levine, and Joshua B Tenenbaum. 2023 · 2023
Cited alongside, same era.
Starcoder: may the source be with you!
Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, et al · 2023
Qiwei Peng, Yekun Chai, and Xuhong Li. 2024 · 2024
Later among the works it cites.
Efficient and Green Large Language Models for Software Engineering: Vision and the Road Ahead
Jieke Shi, Zhou Yang, and David Lo. 2024a · 2024
Later among the works it cites.
LLM4VV: Exploring LLM-as-a-Judge for Validation and Verification Testsuites. In SC24-W: Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 1885–1893
Zachariah Sollenberger, Jay Patel, Christian Munley, Aaron Jarmusch, and Sunita Chandrasekaran. 2024 · 2024
Later among the works it cites.
Source code summarization in the era of large language models
Weisong Sun, Yun Miao, Yuekang Li, Hongyu Zhang, Chunrong Fang, Yi Liu, Gelei Deng, Yang Liu, and Zhenyu Chen. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
G-eval: Nlg evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Cited alongside, same era.
OpenAI. 2023b · 2023
Cited alongside, same era.
Understanding the effectiveness of large language models in code translation
Rangeet Pan, Ali Reza Ibrahimzada, Rahul Krishna, Divya Sankar, Lambert Pouguem Wassi, Michele Merler, Boris Sobolev, Raju Pavuluri, Saurabh Sinha, and Reyhaneh Jabbarvand. 2023 · 2023
Cited alongside, same era.
Compressing Pre-trained Models of Code into 3 MB. In Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering (Rochester, MI, USA) (ASE ’22) . Association for Computing Machinery, New York, NY, USA, Article 24, 12 pages
Jieke Shi, Zhou Yang, Bowen Xu, Hong Jin Kang, and David Lo. 2023 · 2023
Cited alongside, same era.
BatchEval: Towards Human-like Text Evaluation
Peiwen Yuan, Shaoxiong Feng, Yiwei Li, Xinglin Wang, Boyuan Pan, Heda Wang, and Kan Li. 2023 · 2023
Cited alongside, same era.
Codegeex: A pre-trained model for code generation with multilingual benchmarking on humaneval-x. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 5673–5684
Qinkai Zheng, Xiao Xia, Xu Zou, Yuxiao Dong, Shan Wang, Yufei Xue, Lei Shen, Zihan Wang, Andi Wang, Yang Li, et al · 2023
Cited alongside, same era.
Codebertscore: Evaluating code generation with pretrained models of code
Shuyan Zhou, Uri Alon, Sumit Agarwal, and Graham Neubig. 2023 · 2023
Cited alongside, same era.
ICE-Score: Instructing Large Language Models to Evaluate Code
Terry Yue Zhuo. 2023 · 2023
Cited alongside, same era.
Sijun Tan, Siyuan Zhuang, Kyle Montgomery, William Y Tang, Alejandro Cuadron, Chenguang Wang, Raluca Ada Popa, and Ion Stoica. 2024 · 2024
Later among the works it cites.
CodeJudge: Evaluating Code Generation with Large Language Models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, Miami, Florida, USA, 20032–20051
Weixi Tong and Tianyi Zhang. 2024 · 2024
Later among the works it cites.
Software Testing With Large Language Models: Survey, Landscape, and Vision
Junjie Wang, Yuchao Huang, Chunyang Chen, Zhe Liu, Song Wang, and Qing Wang. 2024 · 2024
Later among the works it cites.
Martin Weyssow, Aton Kamanda, Xin Zhou, and Houari Sahraoui. 2024 · 2024
Later among the works it cites.
Can Large Language Models Serve as Evaluators for Code Summarization?
Yang Wu, Yao Wan, Zhaoyang Chu, Wenting Zhao, Ye Liu, Hongyu Zhang, Xuanhua Shi, and Philip S Yu. 2024 · 2024
Later among the works it cites.
Efficient Adversarial Training in LLMs with Continuous Attacks
Sophie Xhonneux, Alessandro Sordoni, Stephan Günnemann, Gauthier Gidel, and Leo Schwinn. 2024 · 2024
Later among the works it cites.
Human-Like Code Quality Evaluation through LLM-based Recursive Semantic Comprehension
Fangzhou Xu, Sai Zhang, Zhenchang Xing, Xiaowang Zhang, Yahong Han, and Zhiyong Feng. 2024 · 2024
Later among the works it cites.
Stealthy Backdoor Attack for Code Models
Zhou Yang, Bowen Xu, Jie M. Zhang, Hong Jin Kang, Jieke Shi, Junda He, and David Lo. 2024 · 2024
Later among the works it cites.
Justice or prejudice? quantifying biases in llm-as-a-judge
Jiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen, Qihui Zhang, Nuno Moniz, Tian Gao, Werner Geyer, Chao Huang, Pin-Yu Chen, et al · 2024
Later among the works it cites.
Transagent: An llm-based multi-agent system for code translation
Zhiqiang Yuan, Weitong Chen, Hanlin Wang, Kai Yu, Xin Peng, and Yiling Lou. 2024 · 2024
Later among the works it cites.
Commit0: Library Generation from Scratch
Wenting Zhao, Nan Jiang, Celine Lee, Justin T Chiu, Claire Cardie, Matthias Gallé, and Alexander M Rush. 2024 · 2024
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2024
Later among the works it cites.
Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions
Terry Yue Zhuo, Minh Chien Vu, Jenny Chim, Han Hu, Wenhao Yu, Ratnadira Widyasari, Imam Nur Bani Yusuf, Haolan Zhan, Junda He, Indraneil Paul, et al · 2024
Later among the works it cites.
Cursor - The AI Code Editor — cursor.com
Anysphere. 2025 · 2025
Closest in time.
The Current Challenges of Software Engineering in the Era of Large Language Models
Cuiyun Gao, Xing Hu, Shan Gao, Xin Xia, and Zhi Jin. 2025 · 2025
Closest in time.
GitHub Copilot: Your AI Pair Programmer
GitHub. 2025 · 2025
Closest in time.
GitHub Copilot · Your AI pair programmer
GitHub. 2025 · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
How to evaluate the cognitive abilities of LLMs
Anna A Ivanova. 2025 · 2025
Closest in time.
Advancing Reasoning in Large Language Models: Promising Methods and Approaches
Avinash Patil. 2025 · 2025
Closest in time.
Confront Insider Threat: Precise Anomaly Detection in Behavior Logs Based on LLM Fine-Tuning. In Proceedings of the 31st International Conference on Computational Linguistics , Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al-Khalifa, Barbara Di Eugenio, and Steven Schockaert (Eds.). Association for Computational Linguistics, Abu Dhabi, UAE, 8589–8601
Shuang Song, Yifei Zhang, and Neng Gao. 2025 · 2025
Closest in time.
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
Ruiqi Wang, Jiyu Guo, Cuiyun Gao, Guodong Fan, Chun Yong Chong, and Xin Xia. 2025 · 2025
Closest in time.
Large Language Model Critics for Execution-Free Evaluation of Code Changes
Aashish Yadavally, Hoan Nguyen, Laurent Callot, and Gauthier Guinet. 2025 · 2025
Closest in time.
CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?. In Proceedings of the 31st International Conference on Computational Linguistics , Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al-Khalifa, Barbara Di Eugenio, and Steven Schockaert (Eds.). Association for Computational Linguistics, Abu Dhabi, UAE, 73–95
Yuwei Zhao, Ziyang Luo, Yuchen Tian, Hongzhan Lin, Weixiang Yan, Annan Li, and Jing Ma. 2025 · 2025
Closest in time.