Fetching the paper…
Reading the bibliography…
Coding agents powered by large language models have shown impressive capabilities in software engineering tasks, but evaluating their performance across diverse programming languages and real-world scenarios remains challenging.
Toward automatic program synthesis
Zohar Manna and Richard J. Waldinger · 1971
Earlier work this paper cites.
Monetary relationships: a view from Threadneedle Street
Charles Goodhart · 1976
Earlier work this paper cites.
Program synthesis
Sumit Gulwani, Oleksandr Polozov, Rishabh Singh, et al · 2017
Earlier work this paper cites.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Measuring coding challenge competence with APPS
Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
Mconala: a benchmark for code generation from multiple natural languages
Zhiruo Wang, Grace Cuenca, Shuyan Zhou, Frank F Xu, and Graham Neubig · 2022
Earlier work this paper cites.
Cert: continual pre-training on sketches for library-oriented code generation
Daoguang Zan, Bei Chen, Dejian Yang, Zeqi Lin, Minsu Kim, Bei Guan, Yongji Wang, Weizhu Chen, and Jian-Guang Lou · 2022
Earlier work this paper cites.
Multi-lingual evaluation of code generation models
Ben Athiwaratkun, Sanjay Krishna Gouda, Zijian Wang, Xiaopeng Li, Yuchen Tian, Ming Tan, Wasi Uddin Ahmad, Shiqi Wang, Qing Sun, Mingyue Shang, Sujan Kumar Gonugondla, Hantian Ding, Varun Kumar, Nathan Fulton, Arash Farahani, Siddhartha Jain, Robert Giaquinto, Haifeng Qian, Murali Krishna Ramanathan, Ramesh Nallapati, Baishakhi Ray, Parminder Bhatia, Sudipta Sengupta, Dan Roth, and Bing Xiang · 2023
Earlier work this paper cites.
Benchmark probing: Investigating data leakage in large language models
Chunyuan Deng, Yilun Zhao, Xiangru Tang, Mark Gerstein, and Arman Cohan · 2023
Earlier work this paper cites.
Longcoder: A long-range pre-trained language model for code completion
Daya Guo, Canwen Xu, Nan Duan, Jian Yin, and Julian McAuley · 2023
Cited alongside, same era.
Codegen: An open large language model for code with multi-turn program synthesis
Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong · 2023
Cited alongside, same era.
Stack overflow developer survey 2023
Stack Overflow · 2023
Cited alongside, same era.
Code translation with compiler representations
Marc Szafraniec, Baptiste Roziere, Hugh James Leather, Patrick Labatut, Francois Charton, and Gabriel Synnaeve · 2023
Cited alongside, same era.
Can llms replace manual annotation of software engineering artifacts?
Toufique Ahmed, Premkumar Devanbu, Christoph Treude, and Michael Pradel · 2024
Cited alongside, same era.
SWE-bench: Can language models resolve real-world github issues?
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R Narasimhan · 2024
Later among the works it cites.
Repoagent: An llm-powered open-source framework for repository-level code documentation generation, 2024
Qinyu Luo, Yining Ye, Shihao Liang, Zhong Zhang, Yujia Qin, Yaxi Lu, Yesai Wu, Xin Cong, Yankai Lin, Yingli Zhang, Xiaoyin Che, Zhiyuan Liu, and Maosong Sun · 2024
Later among the works it cites.
DebugBench: Evaluating debugging capability of large language models
Runchu Tian, Yining Ye, Yujia Qin, Xin Cong, Yankai Lin, Yinxu Pan, Yesai Wu, Hui Haotian, Liu Weichuan, Zhiyuan Liu, and Maosong Sun · 2024
Later among the works it cites.
Agentless: Demystifying llm-based software engineering agents, 2024
Chunqiu Steven Xia, Yinlin Deng, Soren Dunn, and Lingming Zhang · 2024
Later among the works it cites.
Swe-bench-java: A github issue resolving benchmark for java, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nadia Alshahwan, Jubin Chheda, Anastasia Finogenova, Beliz Gokkaya, Mark Harman, Inna Harper, Alexandru Marginean, Shubho Sengupta, and Eddy Wang · 2024
Cited alongside, same era.
CoCoMIC: Code completion by jointly modeling in-file and cross-file context
Yangruibo Ding, Zijian Wang, Wasi Ahmad, Murali Krishna Ramanathan, Ramesh Nallapati, Parminder Bhatia, Dan Roth, and Bing Xiang · 2024
Cited alongside, same era.
Aider is ai pair programming in your terminal
Paul Gauthier · 2024
Cited alongside, same era.
A survey on large language models for code generation
Juyong Jiang, Fan Wang, Jiasi Shen, Sungju Kim, and Sunghun Kim · 2024
Cited alongside, same era.
Swe-agent: Agent-computer interfaces enable automated software engineering
John Yang, Carlos E Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press
Cited in the paper.
Swe-bench multimodal: Do ai systems generalize to visual software domains?, 2024b
John Yang, Carlos E. Jimenez, Alex L. Zhang, Kilian Lieret, Joyce Yang, Xindi Wu, Ori Press, Niklas Muennighoff, Gabriel Synnaeve, Karthik R. Narasimhan, Diyi Yang, Sida I. Wang, and Ofir Press
Cited in the paper.
Daoguang Zan, Zhirong Huang, Ailun Yu, Shaoxin Lin, Yifan Shi, Wei Liu, Dong Chen, Zongshuai Qi, Hao Yu, Lei Yu, Dezhi Ran, Muhan Zeng, Bo Shen, Pan Bian, Guangtai Liang, Bei Guan, Pengjie Huang, Tao Xie, Yongji Wang, and Qianxiang Wang · 2024
Later among the works it cites.
Agent-as-a-judge: Evaluate agents with agents
Mingchen Zhuge, Changsheng Zhao, Dylan Ashley, Wenyi Wang, Dmitrii Khizbullin, Yunyang Xiong, Zechun Liu, Ernie Chang, Raghuraman Krishnamoorthi, Yuandong Tian, et al · 2024
Later among the works it cites.
Introducing SWE-bench Verified, 2024
Neil Chowdhury, James Aung, Chan Jun Shern, Oliver Jaffe, Dane Sherburn, Giulio Starace, Evan Mays, Rachel Dias, Marwan Aljubeh, Mia Glaese, Carlos E. Jimenez, John Yang, Leyton Ho, Tejal Patwardhan, Kevin Liu, and Aleksander Madry · 2025
Closest in time.
Swe-lancer: Can frontier llms earn 1 million from real-world freelance software engineering?
Samuel Miserendino, Michele Wang, Tejal Patwardhan, and Johannes Heidecke · 2025
Closest in time.
Baxbench: Can llms generate correct and secure backends?
Mark Vero, Niels Mündler, Victor Chibotaru, Veselin Raychev, Maximilian Baader, Nikola Jovanović, Jingxuan He, and Martin Vechev · 2025
Closest in time.