Fetching the paper…
Reading the bibliography…
Code generation agents powered by large language models (LLMs) are revolutionizing the software development paradigm.
J. Herrington, Code generation in action . Manning Publications Co., 2003
2003
Earlier work this paper cites.
E. Kitzelmann, “Inductive programming: A survey of program synthesis techniques,” in International Workshop on Approaches and Applications of Inductive Programming (AAIP) , 2009, pp. 50–73
2009
Earlier work this paper cites.
G. Fraser and A. Arcuri, “Evosuite: On the challenges of test case generation in the real world,” in ICST . IEEE Computer Society, 2013, pp. 362–369
2013
Earlier work this paper cites.
W. Ling, E. Grefenstette, K. M. Hermann, T. Kočiskỳ, A. Senior, F. Wang, and P. Blunsom, “Latent predictor networks for code generation,” in Meeting of the Association for Computational Linguistics (ACL) , 2016, pp. 599–609
2016
Earlier work this paper cites.
P. Yin and G. Neubig, “A syntactic neural model for general-purpose code generation,” in Meeting of the Association for Computational Linguistics (ACL) , 2017, pp. 440–450
2017
Earlier work this paper cites.
M. Rabinovich, M. Stern, and D. Klein, “Abstract syntax networks for code generation and semantic parsing,” in Meeting of the Association for Computational Linguistics (ACL) , 2017, pp. 1139–1149
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Conference on Neural Information Processing Systems (NeurIPS) , 2017
2017
Earlier work this paper cites.
A. Svyatkovskiy, S. K. Deng, S. Fu, and N. Sundaresan, “Intellicode compose: Code generation using transformer,” in ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering , 2020, pp. 1433–1443
2020
Earlier work this paper cites.
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel et al. , “Retrieval-augmented generation for knowledge-intensive nlp tasks,” in Conferecce on Neural Information Processing Systems (NeurIPS) , 2020, pp. 9459–9474
2020
Earlier work this paper cites.
Y. Wang, W. Wang, S. Joty, and S. C. Hoi, “CodeT5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation,” in Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2021, pp. 8696–8708
2021
Earlier work this paper cites.
S. Lu, D. Guo, S. Ren, J. Huang, A. Svyatkovskiy, A. Blanco, C. Clement, D. Drain, D. Jiang, D. Tang et al. , “CodeXGLUE: A machine learning benchmark dataset for code understanding and generation,” in Conference on Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track , 2021
2021
Earlier work this paper cites.
D. Guo, S. Ren, S. Lu, Z. Feng, D. Tang, S. Liu, L. Zhou, N. Duan, A. Svyatkovskiy, S. Fu et al. , “GraphCodeBERT: Pre-training code representations with data flow,” in International Conference on Learning Representations , 2021
2021
Earlier work this paper cites.
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman et al. , “Evaluating large language models trained on code,” 2021
2021
Earlier work this paper cites.
J. Austin, A. Odena, M. I. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. J. Cai, M. Terry, Q. V. Le, and C. Sutton, “Program synthesis with large language models,” 2021
2021
Earlier work this paper cites.
D. Hendrycks, S. Basart, S. Kadavath, M. Mazeika, A. Arora, E. Guo, C. Burns, S. Puranik, H. He, D. Song, and J. Steinhardt, “Measuring coding challenge competence with APPS,” in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , “Training language models to follow instructions with human feedback,” in Conference on Neural Information Processing Systems (NeurIPS) , 2022, pp. 27 730–27 744
2022
Earlier work this paper cites.
Y. Li, D. Choi, J. Chung, N. Kushman, J. Schrittwieser, R. Leblond, T. Eccles, J. Keeling, F. Gimeno, A. Dal Lago et al. , “Competition-level code generation with Alphacode,” Science , vol. 378, no. 6624, pp. 1092–1097, 2022
2022
Earlier work this paper cites.
D. Guo, S. Lu, N. Duan, Y. Wang, M. Zhou, and J. Yin, “UniXcoder: Unified cross-modal pre-training for code representation,” pp. 7212–7225, 2022
2022
Earlier work this paper cites.
F. F. Xu, B. Vasilescu, and G. Neubig, “In-IDE code generation from natural language: Promise and challenges,” ACM Transactions on Software Engineering and Methodology (TOSEM) , vol. 31, no. 2, pp. 1–47, 2022
2022
Earlier work this paper cites.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou et al. , “Chain-of-thought prompting elicits reasoning in large language models,” in Conference on Neural Information Processing Systems (NeurIPS) , 2022, pp. 24 824–24 837
2022
Earlier work this paper cites.
R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, X. Jiang, K. Cobbe, T. Eloundou, G. Krueger, K. Button, M. Knight, B. Chess, and J. Schulman, “WebGPT: Browser-assisted question-answering with human feedback,” 2022
2022
Earlier work this paper cites.
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar et al. , “LLaMA: Open and efficient foundation language models,” 2023
2023
Earlier work this paper cites.
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale et al. , “LLaMA 2: Open foundation and fine-tuned chat models,” 2023
2023
Earlier work this paper cites.
B. Roziere, J. Gehring, F. Gloeckle, S. Sootla, I. Gat, X. E. Tan, Y. Adi, J. Liu, R. Sauvestre, T. Remez et al. , “Code LLaMA: Open foundation models for code,” 2023
2023
Earlier work this paper cites.
H. Le, H. Chen, A. Saha, A. Gokul, D. Sahoo, and S. Joty, “CodeChain: Towards modular code generation through chain of self-revisions with representative sub-modules,” in International Conference on Learning Representations (ICLR) , 2023
2023
Earlier work this paper cites.
C. Qian, W. Liu, H. Liu, N. Chen, Y. Dang, J. Li, C. Yang, W. Chen, Y. Su, X. Cong et al. , “Chatdev: Communicative agents for software development,” in Meeting of the Association for Computational Linguistics (ACL) , 2023, pp. 15 174–15 186
2023
Earlier work this paper cites.
S. Hong, X. Zheng, J. Chen, Y. Cheng, J. Wang, C. Zhang, Z. Wang, S. K. S. Yau, Z. Lin, L. Zhou et al. , “Metagpt: Meta programming for multi-agent collaborative framework,” in International Conference on Learning Representations (ICLR) , 2023
2023
Earlier work this paper cites.
F. Mu, L. Shi, S. Wang, Z. Yu, B. Zhang, C. Wang, S. Liu, and Q. Wang, “ClarifyGPT: Empowering LLM-based code generation with intention clarification,” 2023
2023
Earlier work this paper cites.
M. Schäfer, S. Nadi, A. Eghbali, and F. Tip, “Adaptive test generation using a large language model,” 2023
2023
Earlier work this paper cites.
B. A. Becker, P. Denny, J. Finnie-Ansley, A. Luxton-Reilly, J. Prather, and E. A. Santos, “Programming is hard-or at least it used to be: Educational opportunities and challenges of AI code generation,” in ACM Technical Symposium on Computer Science Education V. 1 (SIGCSE) , 2023, pp. 500–506
2023
Earlier work this paper cites.
E. Kasneci, K. Seßler, S. Küchemann, M. Bannert, D. Dementieva, F. Fischer, U. Gasser, G. Groh, S. Günnemann, E. Hüllermeier et al. , “ChatGPT for good? on opportunities and challenges of large language models for education,” Learning and Individual Differences , vol. 103, p. 102274, 2023
2023
Earlier work this paper cites.
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong et al. , “A survey of large language models,” 2023
2023
Earlier work this paper cites.
H. Naveed, A. U. Khan, S. Qiu, M. Saqib, S. Anwar, M. Usman, N. Akhtar, N. Barnes, and A. Mian, “A comprehensive overview of large language models,” ACM Transactions on Intelligent Systems and Technology , 2023
2023
Earlier work this paper cites.
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom, “Toolformer: Language models can teach themselves to use tools,” in Conference on Neural Information Processing Systems (NeurIPS) , 2023, pp. 68 539–68 551
2023
Earlier work this paper cites.
L. Gao, A. Madaan, S. Zhou, U. Alon, P. Liu, Y. Yang, J. Callan, and G. Neubig, “Pal: Program-aided language models,” in International Conference on Machine Learning (ICML) , 2023
2023
Earlier work this paper cites.
J. S. Park, J. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein, “Generative agents: Interactive simulacra of human behavior,” in Proceedings of the 36th annual acm symposium on user interface software and technology , 2023, pp. 1–22
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
L. Giray, “Prompt engineering with ChatGPT: a guide for academic writers,” Annals of biomedical engineering , vol. 51, no. 12, pp. 2629–2633, 2023
2023
Earlier work this paper cites.
J. White, Q. Fu, S. Hays, M. Sandborn, C. Olea, H. Gilbert, A. Elnashar, J. Spencer-Smith, and D. C. Schmidt, “A prompt pattern catalog to enhance prompt engineering with chatgpt,” 2023
2023
Earlier work this paper cites.
Y. Gao, Y. Xiong, X. Gao, K. Jia, J. Pan, Y. Bi, Y. Dai, J. Sun, H. Wang, and H. Wang, “Retrieval-augmented generation for large language models: A survey,” 2023
2023
Earlier work this paper cites.
J. Li, C. Tao, J. Li, G. Li, Z. Jin, H. Zhang, Z. Fang, and F. Liu, “Large language model-aware in-context learning for code generation,” ACM Transactions on Software Engineering and Methodology , 2023
2023
Earlier work this paper cites.
J. Wei, J. Wei, Y. Tay, D. Tran, A. Webson, Y. Lu, X. Chen, H. Liu, D. Huang, D. Zhou et al. , “Larger language models do in-context learning differently,” 2023
2023
Earlier work this paper cites.
D. Huang, J. M. Zhang, M. Luck, Q. Bu, Y. Qing, and H. Cui, “AgentCoder: Multi-agent-based code generation with iterative testing and optimisation,” 2023
2023
Earlier work this paper cites.
K. Zhang, H. Zhang, G. Li, J. Li, Z. Li, and Z. Jin, “ToolCoder: Teach code generation models to use API search tools,” 2023
2023
Earlier work this paper cites.
A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang et al. , “Self-refine: Iterative refinement with self-feedback,” in Conference on Neural Information Processing Systems (NeurIPS) , 2023, pp. 46 534–46 594
2023
Earlier work this paper cites.
K. Zhang, Z. Li, J. Li, G. Li, and Z. Jin, “Self-Edit: Fault-aware code editor for code generation,” in Meeting of the Association for Computational Linguistics (ACL) , 2023, pp. 769–787
2023
Earlier work this paper cites.
T. Chang, S. Chen, G. Fan, and Z. Feng, “A self-iteration code generation method based on large language models,” in International Conference on Parallel and Distributed Systems (ICPADS) , 2023, pp. 275–281
2023
Earlier work this paper cites.
X. Chen, M. Lin, N. Schärli, and D. Zhou, “Teaching large language models to self-debug,” in Meeting Of The Association For Computational Linguistics (ACL) , 2023
2023
Earlier work this paper cites.
S. Holt, M. R. Luyten, and M. van der Schaar, “L2MAC: Large language model automatic computer for extensive code generation,” 2023
2023
Earlier work this paper cites.
D. Chen, H. Wang, Y. Huo, Y. Li, and H. Zhang, “Gamegpt: Multi-agent collaborative framework for game development,” 2023
2023
Earlier work this paper cites.
X. Jiang, Y. Dong, L. Wang, Z. Fang, Q. Shang, G. Li, Z. Jin, and W. Jiao, “Self-planning code generation with large language models,” ACM Transactions on Software Engineering and Methodology , vol. 33, no. 7, pp. 1–30, 2024
2024
Earlier work this paper cites.
Z. Rasheed, M. Waseem, M. Saari, K. Systä, and P. Abrahamsson, “Codepori: Large scale model for autonomous software development by using multi-agents,” 2024
2024
Earlier work this paper cites.
S. Manish, “An autonomous multi-agent llm framework for agile software development,” International Journal of Trend in Scientific Research and Development , vol. 8, no. 5, pp. 892–898, 2024
2024
Earlier work this paper cites.
Y. Dong, X. Jiang, Z. Jin, and G. Li, “Self-collaboration code generation via ChatGPT,” ACM Transactions on Software Engineering and Methodology , vol. 33, no. 7, pp. 1–38, 2024
2024
Earlier work this paper cites.
Z. Wang, W. Wang, Z. Li, L. Wang, C. Yi, X. Xu, L. Cao, H. Su, S. Chen, and J. Zhou, “Xuat-copilot: Multi-agent collaborative system for automated user acceptance testing with large language model,” 2024
2024
Cited alongside, same era.
N. Baumgartner, P. Iyenghar, T. Schoemaker, and E. Pulvermüller, “Ai-driven refactoring: A pipeline for identifying and correcting data clumps in git repositories,” Electronics , vol. 13, no. 9, p. 1644, 2024
2024
Cited alongside, same era.
Y. Ishibashi and Y. Nishimura, “Self-organized agents: A LLM multi-agent framework toward ultra large-scale code generation and optimization,” 2024
2024
Cited alongside, same era.
K. Zhang, J. Li, G. Li, X. Shi, and Z. Jin, “CodeAgent: Enhancing code generation with tool-integrated agent systems for real-world repo-level coding challenges,” in Meeting of the Association for Computational Linguistics (ACL) , 2024
2024
Cited alongside, same era.
R. Wang, X. Han, L. Ji, S. Wang, T. Baldwin, and H. Li, “Toolgen: Unified tool retrieval and calling via generation,” in International Conference on Learning Representations (ICLR) , 2025
2025
Closest in time.
R. Sapkota, K. I. Roumeliotis, and M. Karkee, “Vibe coding vs. agentic coding: Fundamentals and practical implications of agentic AI,” 2025
2025
Closest in time.
Q. Xu, G. Wang, L. Briand, and K. Liu, “A multi-agent llm-based juit test generation with strong oracles,” 2025
2025
Closest in time.
M. T. Dearing, Y. Tao, X. Wu, Z. Lan, and V. Taylor, “Leveraging llms to automate energy-aware refactoring of parallel scientific codes,” 2025
2025
Closest in time.
H. Peng, A. Gupte, R. Hasler, N. J. Eliopoulos, C.-C. Ho, R. Mantri, L. Deng, K. Läufer, G. K. Thiruvathukal, and J. C. Davis, “Sysllmatic: Large language models are software system optimizers,” 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
W. Shi, Y. Zhang, X. Xing, and J. Xu, “Harnessing large language models for seed generation in greybox fuzzing,” 2024
2024
Cited alongside, same era.
H. Jin, L. Huang, H. Cai, J. Yan, B. Li, and H. Chen, “From LLMs to LLM-based agents for software engineering: A survey of current, challenges and future,” 2024
2024
Cited alongside, same era.
J. Liu, K. Wang, Y. Chen, X. Peng, Z. Chen, L. Zhang, and Y. Lou, “Large language model-based agents for software engineering: A survey,” 2024
2024
Cited alongside, same era.
Y. Wang, W. Zhong, Y. Huang, E. Shi, M. Yang, J. Chen, H. Li, Y. Ma, Q. Wang, and Z. Zheng, “Agents in software engineering: Survey, landscape, and vision,” 2024
2024
Cited alongside, same era.
Y. Chang, X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y. Wang et al. , “A survey on evaluation of large language models,” ACM transactions on intelligent systems and technology , vol. 15, no. 3, pp. 1–45, 2024
2024
Cited alongside, same era.
D. Guo, Q. Zhu, D. Yang, Z. Xie, K. Dong, W. Zhang, G. Chen, X. Bi, Y. Wu, Y. Li et al. , “Deepseek-coder: When the large language model meets programming–the rise of code intelligence,” 2024
2024
Cited alongside, same era.
B. Hui, J. Yang, Z. Cui, J. Yang, D. Liu, L. Zhang, T. Liu, J. Zhang, B. Yu, K. Lu et al. , “Qwen2. 5-coder technical report,” 2024
2024
Cited alongside, same era.
L. Wang, C. Ma, X. Feng, Z. Zhang, H. Yang, J. Zhang, Z. Chen, J. Tang, X. Chen, Y. Lin et al. , “A survey on large language model based autonomous agents,” Frontiers of Computer Science , vol. 18, no. 6, p. 186345, 2024
2024
Cited alongside, same era.
2025
Closest in time.
C. Foster, A. Gulati, M. Harman, I. Harper, K. Mao, J. Ritchey, H. Robert, and S. Sengupta, “Mutation-guided llm-based test generation at meta,” 2025
2025
Closest in time.
J. He, C. Treude, and D. Lo, “Llm-based multi-agent systems for software engineering: Literature review, vision, and the road ahead,” ACM Transactions on Software Engineering and Methodology , vol. 34, no. 5, pp. 1–30, 2025
2025
Closest in time.
N. Huynh and B. Lin, “Large language models for code generation: A comprehensive survey of challenges, techniques, evaluation, and applications,” 2025
2025
Closest in time.
E. Wang, F. Cassano, C. Wu, Y. Bai, W. Song, V. Nath, Z. Han, S. Hendryx, S. Yue, and H. Zhang, “Planning in natural language improves LLM search for code generation,” in International Conference on Learning Representations (ICLR) , 2025
2025
Closest in time.
Z. Xi, W. Chen, X. Guo, W. He, Y. Ding, B. Hong, M. Zhang, J. Wang, S. Jin, E. Zhou et al. , “The rise and potential of large language model based agents: A survey,” Science China Information Sciences , vol. 68, no. 2, p. 121101, 2025
2025
Closest in time.
J. Li, H. Le, Y. Zhou, C. Xiong, S. Savarese, and D. Sahoo, “Codetree: agent-guided tree search for code generation with large language models,” in Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics (NAACL) , 2025
2025
Closest in time.
Y. Han and C. Lyu, “Multi-stage guided code generation for large language models,” Engineering Applications of Artificial Intelligence , vol. 139, no. PA, p. 109491, 2025
2025
Closest in time.
V. Aggarwal, O. Kamal, A. Japesh, Z. Jin, and B. Schölkopf, “DARS: Dynamic action re-sampling to enhance coding agent performance by adaptive tree traversal,” 2025
2025
Closest in time.
C.-T. Ho, H. Ren, and B. Khailany, “Verilogcoder: Autonomous Verilog coding agents with graph-based planning and abstract syntax tree (ast)-based waveform tracing tool,” in AAAI Conference on Artificial Intelligence (AAAI) , vol. 39, no. 1, 2025, pp. 300–307
2025
Closest in time.
K. Zainullina, A. Golubev, M. Trofimova, S. Polezhaev, I. Badertdinov, D. Litvintseva, S. Karasik, F. Fisin, S. Skvortsov, M. Nekrashevich, A. Shevtsov, and B. Yangel, “Guided search strategies in non-serializable environments with applications to software engineering agents,” in International Conference on Machine Learning (ICML) , 2025
2025
Closest in time.
X. Jiang, Y. Dong, Y. Tao, H. Liu, Z. Jin, W. Jiao, and G. Li, “ROCODE: Integrating backtracking mechanism and program analysis in large language models for code generation,” in IEEE/ACM International Conference on Software Engineering (ICSE) , 2025, pp. 670–670
2025
Closest in time.
Y. Lu, F. Ye, J. Li, Q. Gao, C. Liu, H. Luo, N. Du, X. Li, and F. Ren, “CodeTool: Enhancing programmatic tool invocation of LLMs via process supervision,” 2025
2025
Closest in time.
M. Acharya, Y. Zhang, K. Leach, and Y. Huang, “Optimizing code runtime performance through context-aware retrieval-augmented generation,” in International Conference on Program Comprehension (ICPC) , 2025, pp. 1–5
2025
Closest in time.
M. Athale and V. Vaddina, “Knowledge graph based repository-level code generation,” in IEEE/ACM International Workshop on Large Language Models for Code (LLM4Code) , 2025, pp. 169–176
2025
Closest in time.
Y. Zhang, X. Zhao, Z. Z. Wang, C. Yang, J. Wei, and T. Wu, “cAST: Enhancing code retrieval-augmented generation with structural chunking via abstract syntax tree,” 2025
2025
Closest in time.
Y. Lai, S. Lee, G. Chen, S. Poddar, M. Hu, D. Z. Pan, and P. Luo, “AnalogCoder: Analog circuit design via training-free code generation,” in AAAI Conference on Artificial Intelligence (AAAI) , no. 1, 2025, pp. 379–387
2025
Closest in time.
F. Lin, D. J. Kim et al. , “Soen-101: Code generation by emulating software process models using large language model agents,” in International Conference on Software Engineering (ICSE) , 2025, pp. 1527–1539
2025
Closest in time.
Y. Hu, Q. Zhou, Q. Chen, X. Li, L. Liu, D. Zhang, A. Kachroo, T. Oz, and O. Tripp, “QualityFlow: An agentic workflow for program synthesis controlled by LLM quality checks,” 2025
2025
Closest in time.
A. Rahman, V. Cvetkovic, K. Reece, A. Walters, Y. Hassan, A. Tummeti, B. Torres, D. Cooney, M. Ellis, and D. S. Nikolopoulos, “MACRO: A multi-agent system for optimizing hpc code generation using large language models,” 2025
2025
Closest in time.
S. Liu, J. Fang, H. Zhou, Y. Wang, and Z. Meng, “SEW: Self-evolving agentic workflows for automated code generation,” 2025
2025
Closest in time.
Y. Hu, Y. Cai, Y. Du, X. Zhu, X. Liu, Z. Yu, Y. Hou, S. Tang, and S. Chen, “Self-evolving multi-agent collaboration networks for software development,” in International Conference on Learning Representations (ICLR) , 2025
2025
Closest in time.
Y. Li, J. Li, Q. Wang, M. Yang, H. Kong, and S. Wang, “Cogito, ergo sum: A neurobiologically-inspired cognition-memory-growth system for code generation,” 2025
2025
Closest in time.
X. Guo, X. Wang, Y. Chen, S. Li, C. Han, M. Li, and H. Ji, “Syncmind: Measuring agent out-of-sync recovery in collaborative software engineering,” in International Conference on Machine Learning (ICML) , 2025
2025
Closest in time.
Q. Xu, G. Wang, L. Briand, and K. Liu, “Hallucination to consensus: Multi-agent LLMs for end-to-end test generation with accurate oracles,” 2025
2025
Closest in time.
M. A. Islam, M. E. Ali, and M. R. Parvez, “Codesim: Multi-agent code generation and problem solving through simulation-driven planning and debugging,” in Findings of the Association for Computational Linguistics , 2025
2025
Closest in time.
M. H. Nguyen, T. P. Chau, P. X. Nguyen, and N. D. Bui, “AgileCoder: Dynamic collaborative agents for software development based on agile methodology,” in IEEE/ACM International Conference on AI Foundation Models and Software Engineering (FORGE) , 2025, pp. 156–167
2025
Closest in time.
I. Bouzenia, P. Devanbu, and M. Pradel, “Repairagent: An autonomous, llm-based agent for program repair,” in International Conference on Software Engineering (ICSE) , 2025, pp. 694–694
2025
Closest in time.
J. Cen, J. Liu, Z. Li, and J. Wang, “SQLFixAgent: Towards semantic-accurate text-to-SQL parsing via consistency-enhanced multi-agent collaboration,” in AAAI Conference on Artificial Intelligence (AAAI) , no. 1, 2025, pp. 49–57
2025
Closest in time.
Z. Yu, H. Zhang, Y. Zhao, H. Huang, M. Yao, K. Ding, and J. Zhao, “Orcaloca: An llm agent framework for software issue localization,” in International Conference on Machine Learning (ICML) , 2025
2025
Closest in time.
H. Li, Y. Tang, S. Wang, and W. Guo, “Patchpilot: A stable and cost-efficient agentic patching framework,” 2025
2025
Closest in time.
Y. Ma, Y. Li, Y. Dong, X. Jiang, R. Cao, J. Chen, F. Huang, and B. Li, “Thinking longer, not larger: Enhancing software engineering agents via scaling test-time compute,” 2025
2025
Closest in time.
H. Ye, A. Z. Yang, C. Hu, Y. Wang, T. Zhang, and C. Le Goues, “Adverintent-agent: Adversarial reasoning for repair based on inferred program intent,” Proceedings of the ACM on Software Engineering , vol. 2, no. ISSTA, pp. 1398–1420, 2025
2025
Closest in time.
A. Sohrabizadeh, J. Song, M. Liu, R. Roy, C. Lee, J. Raiman, and B. Catanzaro, “Nemotron-cortexa: Enhancing llm agents for software engineering tasks via improved localization and solution diversity,” in International Conference on Machine Learning (ICML) , 2025
2025
Closest in time.
S. Siddeeq, Z. Rasheed, M. A. Sami, M. Hasan, M. Waseem, J. Rasku, M. Saari, K.-K. Kemell, and P. Abrahamsson, “Distributed approach to haskell based applications refactoring with llms based multi-agent systems,” 2025
2025
Closest in time.
Z. Jiang, D. Schmidt, D. Srikanth, D. Xu, I. Kaplan, D. Jacenko, and Y. Wu, “AIDE: AI-driven exploration in the space of code,” 2025
2025
Closest in time.
H. Jia, R. Morris, H. Ye, F. Sarro, and S. Mechtaev, “Automated repair of ambiguous natural language requirements,” 2025
2025
Closest in time.
S. Vijayvargiya, X. Zhou, A. Yerukola, M. Sap, and G. Neubig, “Interactive agents to overcome ambiguity in software engineering,” 2025
2025
Closest in time.
E. A. González, R. Rothkopf, S. Lerner, and N. Polikarpova, “HILDE: Intentional code generation via human-in-the-loop decoding,” 2025
2025
Closest in time.
N. Jain, K. Han, A. Gu, W. Li, F. Yan, T. Zhang, S. Wang, A. Solar-Lezama, K. Sen, and I. Stoica, “Livecodebench: Holistic and contamination free evaluation of large language models for code,” in International Conference on Learning Representations (ICLR) , 2025
2025
Closest in time.
K. Xu, Y. Mao, X. Guan, and Z. Feng, “Web-bench: A LLM code benchmark based on web standards and frameworks,” 2025
2025
Closest in time.
P. Gauthier, “Gpt code editing benchmarks,” https://aider.chat/docs/benchmarks.html#the-benchmark , 2024, [Accessed 21-01-2025]
2025
Closest in time.
Y. Dong, J. Ding, X. Jiang, G. Li, Z. Li, and Z. Jin, “Codescore: Evaluating code generation by learning code execution,” ACM Trans. Softw. Eng. Methodol. , vol. 34, no. 3, pp. 77:1–77:22, 2025
2025
Closest in time.
J. Hu, W. Zheng, Y. Liu, and Y. Liu, “Optimizing token consumption in llms: A nano surge approach for code reasoning efficiency,” 2025
2025
Closest in time.
C. Zhao, C. Deng, C. Ruan, D. Dai, H. Gao, J. Li, L. Zhang, P. Huang, S. Zhou, S. Ma, W. Liang, Y. He, Y. Wang, Y. Liu, and Y. X. Wei, “Insights into deepseek-v3: Scaling challenges and reflections on hardware for ai architectures,” in International Symposium on Computer Architecture (ISCA) , 2025, pp. 1731–1745
2025
Closest in time.
H. Lee, Z. Zhang, H. Lu, and L. Zhang, “Sec-bench: Automated benchmarking of llm agents on real-world software security tasks,” 2025
2025
Closest in time.
K. Zhang, H. Zhang, G. Li, J. You, J. Li, Y. Zhao, and Z. Jin, “Sealign: Alignment training for software engineering agent,” 2025
2025
Closest in time.
S. Siddeeq, M. Waseem, Z. Rasheed, M. M. Hasan, J. Rasku, M. Saari, H. Terho, K. Makela, K.-K. Kemell, and P. Abrahamsson, “Llm-based multi-agent system for intelligent refactoring of haskell code,” 2025
2025
Closest in time.
Z. Chen and L. Jiang, “Evaluating software development agents: Patch patterns, code quality, and issue complexity in real-world github scenarios,” pp. 657–668, 2025
2025
Closest in time.
R. Zhang, N. Javidnia, N. Sheybani, and F. Koushanfar, “Robust and secure code watermarking for large language models via ML/Crypto codesign,” 2025
2025
Closest in time.