Fetching the paper…
Reading the bibliography…
The task of issue resolving is to modify a codebase to generate a patch that addresses a given issue.
Visualization of test information to assist fault localization
J. A. Jones, M. J. Harrold, and J. Stasko · 2002
Earlier work this paper cites.
On the accuracy of spectrum-based fault localization
R. Abreu, P. Zoeteweij, and A. J. Van Gemund · 2007
Earlier work this paper cites.
Mining source code repositories at massive scale using language modeling
M. Allamanis and C. Sutton · 2013
Earlier work this paper cites.
Probabilistic model for code with decision trees
V. Raychev, P. Bielik, and M. Vechev · 2016
Earlier work this paper cites.
Mapping language to code in programmatic context
S. Iyer, I. Konstas, A. Cheung, and L. Zettlemoyer · 2018
Earlier work this paper cites.
Program synthesis with large language models
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, et al · 2021
Earlier work this paper cites.
Multi-lingual evaluation of code generation models
B. Athiwaratkun, S. K. Gouda, Z. Wang, X. Li, Y. Tian, M. Tan, W. U. Ahmad, S. Wang, Q. Sun, M. Shang, et al · 2022
Earlier work this paper cites.
CERT: continual pre-training on sketches for library-oriented code generation
D. Zan, B. Chen, D. Yang, Z. Lin, M. Kim, B. Guan, Y. Wang, W. Chen, and J. Lou · 2022
Earlier work this paper cites.
Multi-lingual evaluation of code generation models
B. Athiwaratkun, S. K. Gouda, Z. Wang, X. Li, Y. Tian, M. Tan, W. U. Ahmad, S. Wang, Q. Sun, M. Shang, et al · 2023
Earlier work this paper cites.
Swe-bench: Can language models resolve real-world github issues?
C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. Narasimhan · 2023
Earlier work this paper cites.
Execution-based evaluation for open-domain code generation
Z. Wang, S. Zhou, D. Fried, and G. Neubig · 2023
Earlier work this paper cites.
Large language models meet nl2code: A survey
D. Zan, B. Chen, F. Zhang, D. Lu, B. Wu, B. Guan, W. Yongji, and J.-G. Lou · 2023
Earlier work this paper cites.
Repocoder: Repository-level code completion through iterative retrieval and generation
F. Zhang, B. Chen, Y. Zhang, J. Keung, J. Liu, D. Zan, Y. Mao, J.-G. Lou, and W. Chen · 2023
Earlier work this paper cites.
Crosscodeeval: A diverse and multilingual benchmark for cross-file code completion
Y. Ding, Z. Wang, W. Ahmad, H. Ding, M. Tan, N. Jain, M. K. Ramanathan, R. Nallapati, P. Bhatia, D. Roth, et al · 2024
Cited alongside, same era.
A deep dive into large language models for automated bug localization and repair
S. B. Hossain, N. Jiang, Q. Zhou, X. Li, W.-H. Chiang, Y. Lyu, H. Nguyen, and O. Tripp · 2024
Cited alongside, same era.
A. Jaech, A. Kalai, A. Lerer, A. Richardson, A. El-Kishky, A. Low, A. Helyar, A. Madry, A. Beutel, A. Carney, et al · 2024
Cited alongside, same era.
A survey on large language models for code generation, 2024
J. Jiang, F. Wang, J. Shen, S. Kim, and S. Kim · 2024
Cited alongside, same era.
Exploring the integration of large language models in industrial test maintenance processes
Codev: Issue resolving with visual data, 2024
L. Zhang, D. Zan, Q. Yang, Z. Huang, D. Chen, B. Shen, T. Liu, Y. Gong, P. Huang, X. Lu, G. Liang, L. Cui, and Q. Wang · 2024
Later among the works it cites.
Augment swe-bench verified agent, 2025
augment code · 2025
Closest in time.
Envbench: A benchmark for automated environment setup, 2025
A. Eliseeva, A. Kovrigin, I. Kholkin, E. Bogomolov, and Y. Zharov · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Closest in time.
An llm-based agent for reliable docker environment configuration
R. Hu, C. Peng, X. Wang, and C. Gao · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Lemner, L. Wahlgren, G. Gay, N. Mohammadiha, J. Liu, and J. Wennerberg · 2024
Cited alongside, same era.
Swt-bench: Testing and validating real-world bug-fixes with code agents
N. Mündler, M. Müller, J. He, and M. Vechev · 2024
Cited alongside, same era.
Benchmarking automated program repair: An extensive study on both real-world and artificial bugs
Y. Ouyang, J. Yang, and L. Zhang · 2024
Cited alongside, same era.
Gitbug-actions: Building reproducible bug-fix benchmarks with github actions
N. Saavedra, A. Silva, and M. Monperrus · 2024
Cited alongside, same era.
Agentless: Demystifying llm-based software engineering agents
C. S. Xia, Y. Deng, S. Dunn, and L. Zhang · 2024
Cited alongside, same era.
SWE-agent: Agent-computer interfaces enable automated software engineering
J. Yang, C. E. Jimenez, A. Wettig, K. Lieret, S. Yao, K. Narasimhan, and O. Press · 2024
Cited alongside, same era.
Codereval: A benchmark of pragmatic code generation with generative pre-trained models
H. Yu, B. Shen, D. Ran, J. Zhang, Q. Zhang, Y. Ma, G. Liang, Y. Li, Q. Wang, and T. Xie · 2024
Cited alongside, same era.
Envgen: Generating and adapting environments via llms for training embodied agents, 2024
A. Zala, J. Cho, H. Lin, J. Yoon, and M. Bansal · 2024
Cited alongside, same era.
Closest in time.
Large language models (llms) for source code analysis: applications, models and datasets, 2025
H. Jelodar, M. Meymani, and R. Razavi-Far · 2025
Closest in time.
Swe-lancer: Can frontier llms earn $1 million from real-world freelance software engineering?
S. Miserendino, M. Wang, T. Patwardhan, and J. Heidecke · 2025
Closest in time.
Openai o3-mini, 2025
OpenAI · 2025
Closest in time.
Code digital twin: Empowering llms with tacit knowledge for complex software maintenance
X. Peng, C. Wang, M. Liu, Y. Lou, and Y. Wu · 2025
Closest in time.
Paperbench: Evaluating ai’s ability to replicate ai research
G. Starace, O. Jaffe, D. Sherburn, J. Aung, C. J. Shern, L. Maksin, R. Dias, E. Mays, B. Kinsella, W. Thompson, J. Heidecke, M. Glaese, T. Patwardhan, and OpenAI · 2025
Closest in time.
Swe-agent remote execution framework, 2025
SWE-agent · 2025
Closest in time.
An empirical study on leveraging images in automated bug report reproduction
D. Wang, Z. Zhang, S. Feng, W. G. Halfond, and T. Yu · 2025
Closest in time.
SWE-bench multimodal: Do ai systems generalize to visual software domains?
J. Yang, C. E. Jimenez, A. L. Zhang, K. Lieret, J. Yang, X. Wu, O. Press, N. Muennighoff, G. Synnaeve, K. R. Narasimhan, D. Yang, S. I. Wang, and O. Press · 2025
Closest in time.