Fetching the paper…
Reading the bibliography…
Code Agent development is an extremely active research area, where a reliable performance metric is critical for tracking progress and guiding new developments.
Program synthesis with large language models
Austin, J., Odena, A., Nye, M. I., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C. J., Terry, M., Le, Q. V., and Sutton, C · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W · 2021
Earlier work this paper cites.
Measuring coding challenge competence with APPS
Hendrycks, D., Basart, S., Kadavath, S., Mazeika, M., Arora, A., Guo, E., Burns, C., Puranik, S., He, H., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Repobench: Benchmarking repository-level code auto-completion systems
Liu, T., Xu, C., and McAuley, J. J · 2023
Earlier work this paper cites.
Aider is SOTA for both SWE Bench and SWE Bench Lite, Jun 2024
Aider · 2024
Earlier work this paper cites.
Model card addendum: Claude 3.5 haiku and upgraded claude 3.5 sonnet
Anthropic · 2024
Earlier work this paper cites.
You name it, I run it: An LLM agent to execute tests of arbitrary projects
Bouzenia, I. and Pradel, M · 2024
Earlier work this paper cites.
Awesome python applications
Hashemi, M · 2024
Cited alongside, same era.
Competition-level problems are effective LLM evaluators
Huang, Y., Lin, Z., Liu, X., Gong, Y., Lu, S., Lei, F., Liang, Y., Shen, Y., Lin, C., Duan, N., and Chen, W · 2024
Cited alongside, same era.
R2E: turning any github repository into a programming agent environment
Jain, N., Shetty, M., Zhang, T., Han, K., Sen, K., and Stoica, I · 2024
Cited alongside, same era.
R2e: Turning any github repository into a programming agent test environment
Jain, N., Shetty, M., Zhang, T., Han, K., Sen, K., and Stoica, I · 2024
Cited alongside, same era.
Swe-bench: Can language models resolve real-world github issues?
Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K. R · 2024
Cited alongside, same era.
Swt-bench: Testing and validating real-world bug-fixes with code agents
Mündler, N., Mueller, M. N., He, J., and Vechev, M · 2024
Code generation with alphacodium: From prompt engineering to flow engineering
Ridnik, T., Kredo, D., and Friedman, I · 2024
Later among the works it cites.
hugovk/top-pypi-packages: Release 2024.12, December 2024
van Kemenade, H., Paterson, C., Thoma, M., Si, R., and Dollenstein, Z · 2024
Later among the works it cites.
Agentless: Demystifying llm-based software engineering agents
Xia, C. S., Deng, Y., Dunn, S., and Zhang, L · 2024
Later among the works it cites.
SWE-agent: Agent Computer Interfaces Enable Software Engineering Language Models, 2024
Yang, J., Jimenez, C. E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K., and Press, O · 2024
Later among the works it cites.
Autocoderover: Autonomous program improvement
Zhang, Y., Ruan, H., Fan, Z., and Roychoudhury, A · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Opendevin: Code less, make more, 2024
OpenDevin · 2024
Cited alongside, same era.
Repairagent: An autonomous, llm-based agent for program repair
Bouzenia, I., Devanbu, P. T., and Pradel, M
Cited in the paper.
Dypybench: A benchmark of executable python software
Bouzenia, I., Krishan, B. P., and Pradel, M
Cited in the paper.
Livecodebench: Holistic and contamination free evaluation of large language models for code
Jain, N., Han, K., Gu, A., Li, W., Yan, F., Zhang, T., Wang, S., Solar-Lezama, A., Sen, K., and Stoica, I
Cited in the paper.
A survey on large language model based autonomous agents
Wang, L., Ma, C., Feng, X., Zhang, Z., Yang, H., Zhang, J., Chen, Z., Tang, J., Chen, X., Lin, Y., Zhao, W. X., Wei, Z., and Wen, J
Cited in the paper.
OpenHands: An Open Platform for AI Software Developers as Generalist Agents, 2024b
Wang, X., Li, B., Song, Y., Xu, F. F., Tang, X., Zhuge, M., Pan, J., Song, Y., Li, B., Singh, J., Tran, H. H., Li, F., Ma, R., Zheng, M., Qian, B., Shao, Y., Muennighoff, N., Zhang, Y., Hui, B., Lin, J., Brennan, R., Peng, H., Ji, H., and Neubig, G
Cited in the paper.
Later among the works it cites.
Openai model docs
OpenAI · 2025
Closest in time.
Statista market insights
Statista · 2025
Closest in time.