Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have achieved strong performance on code generation, but existing methods still struggle with repository-level code generation under executable validation.
Metrics for Assessing a Software System’s Maintainability. In Proceedings of the Conference on Software Maintenance . 337–344
Paul Oman and Jack Hagemeister. 1992 · 1992
Earlier work this paper cites.
Construction and Testing of Polynomials Predicting Software Maintainability
Paul Oman and Jack Hagemeister. 1994 · 1994
Earlier work this paper cites.
Applying the ISO/IEC 25010 quality models to software product. In European Conference on Software Process Improvement . Springer, 492–503
John Estdale and Elli Georgiadou. 2018 · 2018
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Crosscodeeval: A diverse and multilingual benchmark for cross-file code completion
Yangruibo Ding, Zijian Wang, Wasi Ahmad, Hantian Ding, Ming Tan, Nihal Jain, Murali Krishna Ramanathan, Ramesh Nallapati, Parminder Bhatia, Dan Roth, et al · 2023
Earlier work this paper cites.
Swe-bench: Can language models resolve real-world github issues?
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. 2023 · 2023
Earlier work this paper cites.
Repobench: Benchmarking repository-level code auto-completion systems
Tianyang Liu, Canwen Xu, and Julian McAuley. 2023 · 2023
Earlier work this paper cites.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al · 2023
Earlier work this paper cites.
Repocoder: Repository-level code completion through iterative retrieval and generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing . 2471–2484
Fengji Zhang, Bei Chen, Yue Zhang, Jacky Keung, Jin Liu, Daoguang Zan, Yi Mao, Jian-Guang Lou, and Weizhu Chen. 2023 · 2023
Earlier work this paper cites.
Codeplan: Repository-level coding using llms and planning
Ramakrishna Bairi, Atharv Sonwane, Aditya Kanade, Vageesh D C, Arun Iyer, Suresh Parthasarathy, Sriram Rajamani, Balasubramanyan Ashok, and Shashank Shet. 2024 · 2024
Earlier work this paper cites.
Gitchameleon: Unmasking the version-switching capabilities of code generation models
Nizar Islah, Justine Gehring, Diganta Misra, Eilif Muller, Irina Rish, Terry Yue Zhuo, and Massimo Caccia. 2024 · 2024
Earlier work this paper cites.
A survey on llm-based code generation for low-resource and domain-specific programming languages
Sathvik Joel, Jie Wu, and Fatemeh Fard. 2024 · 2024
Earlier work this paper cites.
Deveval: A manually-annotated code generation benchmark aligned with real-world code repositories. In Findings of the Association for Computational Linguistics: ACL 2024 . 3603–3614
Jia Li, Ge Li, Yunfei Zhao, Yongmin Li, Huanyu Liu, Hao Zhu, Lecheng Wang, Kaibo Liu, Zheng Fang, Lanshen Wang, et al · 2024
Cited alongside, same era.
Repograph: Enhancing ai software engineering with repository-level code graph
Siru Ouyang, Wenhao Yu, Kaixin Ma, Zilin Xiao, Zhihan Zhang, Mengzhao Jia, Jiawei Han, Hongming Zhang, and Dong Yu. 2024 · 2024
Cited alongside, same era.
Versicode: Towards version-controllable code generation
Tongtong Wu, Weigang Wu, Xingyu Wang, Kang Xu, Suyu Ma, Bo Jiang, Ping Yang, Zhenchang Xing, Yuan-Fang Li, and Gholamreza Haffari. 2024b · 2024
Cited alongside, same era.
A comprehensive framework for evaluating api-oriented code generation in large language models
Yixi Wu, Pengfei He, Zehao Wang, Shaowei Wang, Yuan Tian, and Tse-Hsun Chen. 2024a · 2024
Cited alongside, same era.
APIMig: A Project-Level Cross-Multi-Version API Migration Framework Based on Evolution Knowledge Graph. In Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence . 7455–7463
Li Kuang, Qi Xie, HaiYang Yang, Yang Yang, Xiang Wei, HaoYue Kang, and YingJie Xia. 2025 · 2025
Later among the works it cites.
Libevolutioneval: A benchmark and study for version-specific code generation. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) . 6826–6840
Sachit Kuhar, Wasi Ahmad, Zijian Wang, Nihal Jain, Haifeng Qian, Baishakhi Ray, Murali Krishna Ramanathan, Xiaofei Ma, and Anoop Deoras. 2025 · 2025
Later among the works it cites.
On the impacts of contexts on repository-level code generation. In Findings of the Association for Computational Linguistics: NAACL 2025 . 1496–1524
Nam Le Hai, Dung Manh Nguyen, and Nghi DQ Bui. 2025 · 2025
Later among the works it cites.
Fea-bench: A benchmark for evaluating repository-level code generation for feature implementation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Execrepobench: Multi-level executable code completion evaluation
Jian Yang, Jiajun Zhang, Jiaxi Yang, Ke Jin, Lei Zhang, Qiyao Peng, Ken Deng, Yibo Miao, Tianyu Liu, Zeyu Cui, et al · 2024
Cited alongside, same era.
Codereval: A benchmark of pragmatic code generation with generative pre-trained models. In Proceedings of the 46th IEEE/ACM International Conference on Software Engineering . 1–12
Hao Yu, Bo Shen, Dezhi Ran, Jiaxin Zhang, Qi Zhang, Yuchi Ma, Guangtai Liang, Ying Li, Qianxiang Wang, and Tao Xie. 2024 · 2024
Cited alongside, same era.
A systematic literature review on large language models for automated program repair
Quanjun Zhang, Chunrong Fang, Yang Xie, YuXiang Ma, Weisong Sun, Yun Yang, and Zhenyu Chen. 2024 · 2024
Cited alongside, same era.
Codemenv: Benchmarking large language models on code migration. In Findings of the Association for Computational Linguistics: ACL 2025 . 2719–2744
Keyuan Cheng, Xudong Shen, Yihao Yang, TengyueWang TengyueWang, Yang Cao, Muhammad Asif Ali, Hanbin Wang, Lijie Hu, and Di Wang. 2025 · 2025
Cited alongside, same era.
NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents
Jingzhe Ding, Shengda Long, Changxin Pu, Huan Zhou, Hongwan Gao, Xiang Gao, Chao He, Yue Hou, Fei Hu, Zhaojian Li, et al · 2025
Cited alongside, same era.
On the Impacts of Contexts on Repository-Level Code Generation
Nam Le Hai, Dung Manh Nguyen, and Nghi D. Q. Bui. 2025 · 2025
Cited alongside, same era.
Repo2run: Automated building executable environment for code repository at scale
Ruida Hu, Chao Peng, Xinchen Wang, Junjielong Xu, and Cuiyun Gao. 2025b · 2025
Cited alongside, same era.
Assessing and advancing benchmarks for evaluating large language models in software engineering tasks
Xing Hu, Feifei Niu, Junkai Chen, Xin Zhou, Junwei Zhang, Junda He, Xin Xia, and David Lo. 2025a · 2025
Cited alongside, same era.
Wei Li, Xin Zhang, Zhongxin Guo, Shaoguang Mao, Wen Luo, Guangyue Peng, Yangyu Huang, Houfeng Wang, and Scarlett Li. 2025 · 2025
Later among the works it cites.
CodeCoR: An LLM-Based Self-Reflective Multi-Agent Framework for Code Generation
Ruwei Pan, Hongyu Zhang, and Chao Liu. 2025 · 2025
Later among the works it cites.
Large Language Models for Code: A Focused Survey of Ten Recent Studies on Methods, Evaluation, and Robustness
Jishnu Sen. 2025 · 2025
Later among the works it cites.
LLM-Based Code Generation: A Systematic Literature Review With Technical and Demographic Insights
Umama, Kamaluddeen Usman Danyaro, Maged Nasser, Abubakar Zakari, Shamsu Abdullahi, Atika Khanzada, Muhammad Muntasir Yakubu, and Sara Shoaib. 2025 · 2025
Later among the works it cites.
Llms meet library evolution: Evaluating deprecated api usage in llm-based code completion. In 2025 ieee/acm 47th international conference on software engineering (icse) . IEEE, 885–897
Chong Wang, Kaifeng Huang, Jian Zhang, Yebo Feng, Lyuye Zhang, Yang Liu, and Xin Peng. 2025 · 2025
Later among the works it cites.
Code to think, think to code: A survey on code-enhanced reasoning and reasoning-driven code intelligence in llms. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing . 2586–2616
Dayu Yang, Tianyang Liu, Daoan Zhang, Antoine Simoulin, Xiaoyi Liu, Yuwei Cao, Zhaopu Teng, Xin Qian, Grey Yang, Jiebo Luo, et al · 2025
Later among the works it cites.
A survey on large language models for code generation
Juyong Jiang, Fan Wang, Jiasi Shen, Sungju Kim, and Sunghun Kim. 2026 · 2026
Closest in time.
Ruwei Pan, Yakun Zhang, Qingyuan Liang, Yueheng Zhu, Chao Liu, Lu Zhang, and Hongyu Zhang. 2026 · 2026
Closest in time.
Environment-Aware Code Generation: How far are We?
Tongtong Wu, Rongyi Chen, Wenjie Du, Suyu Ma, Guilin Qi, Zhenchang Xing, Shahram Khadivi, Ramesh Periyathambi, and Gholamreza Haffari. 2026 · 2026
Closest in time.