Fetching the paper…
Reading the bibliography…
Software issue resolution is a critical challenge in software engineering and has garnered increasing attention in recent years.
A new measure of rank correlation
M. G. Kendall · 1938
Earlier work this paper cites.
The proof and measurement of association between two things
C. Spearman · 1961
Earlier work this paper cites.
Critical values and probability levels for the Wilcoxon rank sum test and the Wilcoxon signed rank test , volume 1
F. Wilcoxon, S. Katti, R. A. Wilcox, et al · 1963
Earlier work this paper cites.
Can llms replace human evaluators? an empirical study of llm-as-a-judge in software engineering
R. Wang, J. Guo, C. Gao, G. Fan, C. Y. Chong, and X. Xia · 1977
Earlier work this paper cites.
Bagging predictors
L. Breiman · 1996
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Y. Freund and R. E. Schapire · 1997
Earlier work this paper cites.
Aiding program comprehension by static and dynamic feature analysis
T. Eisenbarth, R. Koschke, and D. Simon · 2001
Earlier work this paper cites.
Search-based software engineering
M. Harman and B. F. Jones · 2001
Earlier work this paper cites.
Reformulating software engineering as a search problem
J. Clarke, J. J. Dolado, M. Harman, R. Hierons, B. Jones, M. Lumkin, B. Mitchell, S. Mancoridis, K. Rees, M. Roper, et al · 2003
Earlier work this paper cites.
Is combining classifiers with stacking better than selecting the best one?
S. Džeroski and B. Ženko · 2004
Earlier work this paper cites.
Integrating static and dynamic analysis to improve the comprehension of existing web applications
G. A. Di Lucca and M. Di Penta · 2005
Earlier work this paper cites.
Learning a metric for code readability
R. P. Buse and W. R. Weimer · 2009
Earlier work this paper cites.
Pearson correlation coefficient
I. Cohen, Y. Huang, J. Chen, J. Benesty, J. Benesty, J. Chen, Y. Huang, and I. Cohen · 2009
Earlier work this paper cites.
Ensemble learning
R. Polikar · 2012
Earlier work this paper cites.
Trivial compiler equivalence: A large scale empirical study of a simple, fast and effective equivalent mutant detection technique
M. Papadakis, Y. Jia, M. Harman, and Y. Le Traon · 2015
Earlier work this paper cites.
Assertions are strongly correlated with test suite effectiveness
Y. Zhang and A. Mesbah · 2015
Earlier work this paper cites.
Detecting trivial mutant equivalences via compiler optimisations
M. Kintis, M. Papadakis, Y. Jia, N. Malevris, Y. Le Traon, and M. Harman · 2017
Earlier work this paper cites.
Transforming programs and tests in tandem for fault localization
X. Li and L. Zhang · 2017
Earlier work this paper cites.
Predicting the resilience of obfuscated code against symbolic execution attacks via machine learning
B. Sebastian, C. Christian, and P. Alexander · 2017
Earlier work this paper cites.
Ensemble learning: A survey
O. Sagi and L. Rokach · 2018
Earlier work this paper cites.
Deep learning ensemble for hyperspectral image classification
Y. Chen, Y. Wang, Y. Gu, X. He, P. Ghamisi, and X. Jia · 2019
Cited alongside, same era.
Empirical evaluation of mutation-based test case prioritization techniques
D. Shin, S. Yoo, M. Papadakis, and D.-H. Bae · 2019
Cited alongside, same era.
Codebleu: a method for automatic evaluation of code synthesis
S. Ren, D. Guo, S. Lu, L. Zhou, S. Liu, D. Tang, N. Sundaresan, M. Zhou, A. Blanco, and S. Ma · 2020
Cited alongside, same era.
Selecting fault revealing mutants
T. Titcheu Chekam, M. Papadakis, T. F. Bissyandé, Y. Le Traon, and K. Sen · 2020
Cited alongside, same era.
Gan ensemble for anomaly detection
X. Han, X. Chen, and L.-P. Liu · 2021
Cited alongside, same era.
Predictive mutation analysis via the natural language channel in source code
Swe-bench leaderboard
S. bench Team · 2025
Closest in time.
Unified diff python parsing/metadata extraction library
M. Bordese · 2025
Closest in time.
RepairAgent: An Autonomous, LLM-Based Agent for Program Repair
I. Bouzenia, P. Devanbu, and M. Pradel · 2025
Closest in time.
#1 open-source agent on swe-bench verified by combining claude 3.7 and o1
T. Chen and C. Flaherty · 2025
Closest in time.
Augment swe-bench verified agent
A. Code · 2025
Closest in time.
An analysis tool for python that blurs the line between testing and type systems
CrossHair · 2025
Closest in time.
Gemini 2.5 pro
G. DeepMind · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Kim, J. Jeon, S. Hong, and S. Yoo · 2022
Cited alongside, same era.
Towards generating functionally correct code edits from natural language issue descriptions
S. Fakhoury, S. Chakraborty, M. Musuvathi, and S. K. Lahiri · 2023
Cited alongside, same era.
Automated repair of programs from large language models
Z. Fan, X. Gao, M. Mirchev, A. Roychoudhury, and S. H. Tan · 2023
Cited alongside, same era.
Swe-bench: Can language models resolve real-world github issues?
C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. R. Narasimhan · 2023
Cited alongside, same era.
Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation
J. Liu, C. S. Xia, Y. Wang, and L. Zhang · 2023
Cited alongside, same era.
Lms: Understanding code syntax and semantics for code analysis
W. Ma, S. Liu, Z. Lin, W. Wang, Q. Hu, Y. Liu, C. Zhang, L. Nie, L. Li, and Y. Liu · 2023
Cited alongside, same era.
I. Abdelaziz, K. Basu, M. Agarwal, S. Kumaravel, M. Stallone, R. Panda, Y. Rizk, G. Bhargav, M. Crouse, C. Gunasekara, et al · 2024
Cited alongside, same era.
Closest in time.
Omnigirl: A multilingual and multimodal benchmark for github issue resolution
L. Guo, W. Tao, R. Jiang, Y. Wang, J. Chen, X. Liu, Y. Ma, M. Mao, H. Zhang, and Z. Zheng · 2025
Closest in time.
S*: Test time scaling for code generation
D. Li, S. Cao, C. Cao, X. Li, S. Tan, K. Keutzer, J. Xing, J. E. Gonzalez, and I. Stoica · 2025
Closest in time.
Soen-101: Code generation by emulating software process models using large language model agents
F. Lin, D. J. Kim, and T.-H. P. Chen · 2025
Closest in time.
ToolACE: Winning the points of LLM function calling
W. Liu, X. Huang, X. Zeng, xinlong hao, S. Yu, D. Li, S. Wang, W. Gan, Z. Liu, Y. Yu, Z. WANG, Y. Wang, W. Ning, Y. Hou, B. Wang, C. Wu, W. Xinzhi, Y. Liu, Y. Wang, D. Tang, D. Tu, L. Shang, X. Jiang, R. Tang, D. Lian, Q. Liu, and E. Chen · 2025
Closest in time.
Enhancing llm code generation with ensembles: A similarity-based selection approach
T. Mahmud, B. Duan, C. Pasareanu, and G. Yang · 2025
Closest in time.
Specrover: Code intent extraction via llms
H. Ruan, Y. Zhang, and A. Roychoudhury · 2025
Closest in time.
Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning
C. V. Snell, J. Lee, K. Xu, and A. Kumar · 2025
Closest in time.
Fixing large language models’ specification misunderstanding for better code generation
Z. Tian, J. Chen, and X. Zhang · 2025
Closest in time.
https://github.com/bytedance/trae-agent , 2025
Trae-Agent · 2025
Closest in time.
Demystifying llm-based software engineering agents
C. S. Xia, Y. Deng, S. Dunn, and L. Zhang · 2025
Closest in time.
Sealign: Alignment training for software engineering agent
K. Zhang, H. Zhang, G. Li, J. You, J. Li, Y. Zhao, and Z. Jin · 2025
Closest in time.
A. Örwall · 2025
Closest in time.