Fetching the paper…
Reading the bibliography…
In recent years, AI-based software engineering has progressed from pre-trained models to advanced agentic workflows, with Software Development Agents representing the next major leap.
E. E. Cureton, “Rank-biserial correlation,” Psychometrika , vol. 21, no. 3, pp. 287–290, 1956
1956
Earlier work this paper cites.
J. Kincaid, “Derivation of new readability formulas (automated readability index, fog count and flesch reading ease formula) for navy enlisted personnel,” Chief of Naval Technical Training , 1975
1975
Earlier work this paper cites.
B. W. Yap and C. H. Sim, “Comparisons of various types of normality tests,” Journal of Statistical Computation and Simulation , vol. 81, no. 12, pp. 2141–2155, 2011
2011
Earlier work this paper cites.
M. Tufano, F. Palomba, G. Bavota, R. Oliveto, M. Di Penta, A. De Lucia, and D. Poshyvanyk, “When and why your code starts to smell bad (and whether the smells go away),” IEEE Transactions on Software Engineering , vol. 43, no. 11, pp. 1063–1088, 2017
2017
Earlier work this paper cites.
V. Antinyan, M. Staron, and A. Sandberg, “Evaluating code complexity triggers, use of complexity measures and the influence of code complexity on maintenance time,” Empirical Software Engineering , vol. 22, 12 2017
2017
Earlier work this paper cites.
Y. Fan, X. Xia, D. Lo, and A. E. Hassan, “Chaff from the wheat: Characterizing and determining valid bug reports,” IEEE transactions on software engineering , vol. 46, no. 5, pp. 495–525, 2018
2018
Earlier work this paper cites.
F. Falcão, C. Barbosa, B. Fonseca, A. Garcia, M. Ribeiro, and R. Gheyi, “On relating technical, social factors, and the introduction of bugs,” in 2020 IEEE 27th International Conference on Software Analysis, Evolution and Reengineering (SANER) . IEEE, 2020, pp. 378–388
2020
Earlier work this paper cites.
D. Eleyan, A. Othman, and A. Eleyan, “Enhancing software comments readability using flesch reading ease score,” Information , vol. 11, no. 9, p. 430, 2020
2020
Earlier work this paper cites.
S. Reis, R. Abreu, and L. Cruz, “Fixing vulnerabilities potentially hinders maintainability,” Empirical Software Engineering , vol. 26, no. 6, p. 127, 2021
2021
Earlier work this paper cites.
N. Peitek, S. Apel, C. Parnin, A. Brechmann, and J. Siegmund, “Program comprehension and code complexity metrics: An fmri study,” in 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE) . IEEE, 2021, pp. 524–536
2021
Earlier work this paper cites.
J. Zhou, M. Pacheco, Z. Wan, X. Xia, D. Lo, Y. Wang, and A. E. Hassan, “Finding a needle in a haystack: Automated mining of silent vulnerability fixes,” in 2021 36th IEEE/ACM International Conference on Automated Software Engineering (ASE) . IEEE, 2021, pp. 705–716
2021
Earlier work this paper cites.
T. Sharma and M. Kessentini, “Qscored: A large dataset of code smells and quality metrics,” in 2021 IEEE/ACM 18th international conference on mining software repositories (MSR) . IEEE, 2021, pp. 590–594
2021
Earlier work this paper cites.
H. Pearce, B. Ahmad, B. Tan, B. Dolan-Gavitt, and R. Karri, “Asleep at the keyboard? assessing the security of github copilot’s code contributions,” in 2022 IEEE Symposium on Security and Privacy (SP) . IEEE, 2022, pp. 754–768
2022
Earlier work this paper cites.
W. Xiao, H. He, W. Xu, X. Tan, J. Dong, and M. Zhou, “Recommending good first issues in github oss projects,” in Proceedings of the 44th International Conference on Software Engineering , 2022, pp. 1830–1842
2022
Earlier work this paper cites.
S. Karakatič, A. Miloševič, and T. Heričko, “Software system comparison with semantic source code embeddings,” Empirical Software Engineering , vol. 27, no. 3, p. 70, 2022
2022
Earlier work this paper cites.
N. Nguyen and S. Nadi, “An empirical evaluation of github copilot’s code suggestions,” in Proceedings of the 19th International Conference on Mining Software Repositories , 2022, pp. 1–5
2022
Earlier work this paper cites.
M. L. Siddiq and J. C. Santos, “Securityeval dataset: mining vulnerability examples to evaluate machine learning-based code generation techniques,” in Proceedings of the 1st International Workshop on Mining Software Repositories Applications for Privacy and Security , 2022, pp. 29–33
2022
Earlier work this paper cites.
Z. Li, Y. Yu, T. Wang, Y. Lei, Y. Wang, and H. Wang, “To follow or not to follow: Understanding issue/pull-request templates on github,” IEEE Transactions on Software Engineering , vol. 49, no. 4, pp. 2530–2544, 2022
2022
Cited alongside, same era.
Z. Fan, X. Gao, M. Mirchev, A. Roychoudhury, and S. H. Tan, “Automated repair of programs from large language models,” in 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE) . IEEE, 2023, pp. 1469–1481
2023
Cited alongside, same era.
J. Li, G. Li, Y. Li, and Z. Jin, “Structured chain-of-thought prompting for code generation,” ACM Transactions on Software Engineering and Methodology , 2023
2023
Cited alongside, same era.
X. Hou, Y. Zhao, Y. Liu, Z. Yang, K. Wang, L. Li, X. Luo, D. Lo, J. Grundy, and H. Wang, “Large language models for software engineering: A systematic literature review,” ACM Transactions on Software Engineering and Methodology , 2023
2023
Cited alongside, same era.
M. L. Siddiq, L. Roney, J. Zhang, and J. C. D. S. Santos, “Quality assessment of chatgpt generated code and their use by developers,” in Proceedings of the 21st International Conference on Mining Software Repositories , 2024, pp. 152–156
2024
Closest in time.
Y. Liu, C. Tantithamthavorn, Y. Liu, and L. Li, “On the reliability and explainability of language models for program generation,” ACM Transactions on Software Engineering and Methodology , vol. 33, no. 5, pp. 1–26, 2024
2024
Closest in time.
Z. Chen and L. Jiang, “Promise and peril of collaborative code generation models: Balancing effectiveness and memorization,” in Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering , 2024, pp. 493–505
2024
Closest in time.
Y. Liu, T. Le-Cong, R. Widyasari, C. Tantithamthavorn, L. Li, X.-B. D. Le, and D. Lo, “Refining chatgpt-generated code: Characterizing and mitigating code quality issues,” ACM Transactions on Software Engineering and Methodology , vol. 33, no. 5, pp. 1–26, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Mo, Y. Jiang, W. Zhan, D. Wang, and Z. Li, “A comprehensive study on code clones in automated driving software,” in 2023 38th IEEE/ACM International Conference on Automated Software Engineering (ASE) . IEEE, 2023, pp. 1073–1085
2023
Cited alongside, same era.
D. Yan, Z. Gao, and Z. Liu, “A closer look at different difficulty levels code generation abilities of chatgpt,” in 2023 38th IEEE/ACM International Conference on Automated Software Engineering (ASE) . IEEE, 2023, pp. 1887–1898
2023
Cited alongside, same era.
O. Asare, M. Nagappan, and N. Asokan, “Is github’s copilot as bad as humans at introducing vulnerabilities in code?” Empirical Software Engineering , vol. 28, no. 6, p. 129, 2023
2023
Cited alongside, same era.
C. Tony, M. Mutas, N. E. D. Ferreyra, and R. Scandariato, “Llmseceval: A dataset of natural language prompts for security evaluations,” in 2023 IEEE/ACM 20th International Conference on Mining Software Repositories (MSR) . IEEE, 2023, pp. 588–592
2023
Cited alongside, same era.
L. Wang, X. Tang, Y. He, C. Ren, S. Shi, C. Yan, and Z. Li, “Delving into commit-issue correlation to enhance commit message generation models,” in 2023 38th IEEE/ACM International Conference on Automated Software Engineering (ASE) . IEEE, 2023, pp. 710–722
2023
Cited alongside, same era.
J. He and M. Vechev, “Large language models for code: Security hardening and adversarial testing,” in Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security , 2023, pp. 1865–1879
2023
Cited alongside, same era.
R. Toews, “Agents are the future of ai. where are the startup opportunities?” Forbes, July 2024, accessed: 2024-10-08. [Online]. Available: https://www.forbes.com/sites/robtoews/2024/07/09/agents-are-the-future-of-ai-where-are-the-startup-opportunities/
2024
Cited alongside, same era.
H. Yu, B. Shen, D. Ran, J. Zhang, Q. Zhang, Y. Ma, G. Liang, Y. Li, Q. Wang, and T. Xie, “Codereval: A benchmark of pragmatic code generation with generative pre-trained models,” in Proceedings of the 46th IEEE/ACM International Conference on Software Engineering , 2024, pp. 1–12
2024
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
V. Majdinasab, M. J. Bishop, S. Rasheed, A. Moradidakhel, A. Tahir, and F. Khomh, “Assessing the security of github copilot’s generated code-a targeted replication study,” in 2024 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER) . IEEE, 2024, pp. 435–444
2024
Closest in time.
C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. R. Narasimhan, “SWE-bench: Can language models resolve real-world github issues?” in The Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https://openreview.net/forum?id=VTF8yNQM66
2024
Closest in time.
N. Chowdhury, J. Aung, C. Jun Shern, O. Jaffe, D. Sherburn, G. Starace, E. Mays, R. Dias, M. Aljubeh, M. Glaese, C. E. Jimenez, J. Yang, K. Liu, and A. Madry, “Introducing swe-bench verified: A benchmark of human-validated issue–pull request pairs,” OpenAI, Tech. Rep., August 2024. [Online]. Available: https://openai.com/index/introducing-swe-bench-verified/
2024
Closest in time.
N. Johansson, M. Caporuscio, and T. Olsson, “Mapping source code to software architecture by leveraging large language models,” in European Conference on Software Architecture . Springer, 2024, pp. 133–149
2024
Closest in time.
P. Zhang, Y. Wang, X. Liu, Z. Lu, Y. Yang, Y. Li, L. Chen, Z. Wang, C.-A. Sun, X. Yu et al. , “Assessing effectiveness of test suites: What do we know and what should we do?” ACM Transactions on Software Engineering and Methodology , vol. 33, no. 4, pp. 1–32, 2024
2024
Closest in time.
S. Hamer, M. d’Amorim, and L. Williams, “Just another copy and paste? comparing the security vulnerabilities of chatgpt generated code and stackoverflow answers,” in 2024 IEEE Security and Privacy Workshops (SPW) . IEEE, 2024, pp. 87–94
2024
Closest in time.
T. Zhang, Y. Lu, Y. Yu, X. Mao, Y. Zhang, and Y. Zhao, “How do developers adapt code snippets to their contexts? an empirical study of context-based code snippet adaptations,” IEEE Transactions on Software Engineering , 2024
2024
Closest in time.
J. Liu, C. S. Xia, Y. Wang, and L. Zhang, “Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
A. Velasco, “Beyond accuracy: Evaluating source code capabilities in large language models for software engineering,” in Proceedings of the 2024 IEEE/ACM 46th International Conference on Software Engineering: Companion Proceedings , 2024, pp. 162–164
2024
Closest in time.
A. Kavian, M. M. Pourhashem Kallehbasti, S. Kazemi, E. Firouzi, and M. Ghafari, “Llm security guard for code,” in Proceedings of the 28th International Conference on Evaluation and Assessment in Software Engineering , 2024, pp. 600–603
2024
Closest in time.
M. F. Rabbi, A. I. Champa, M. F. Zibran, and M. R. Islam, “Ai writes, we analyze: The chatgpt python code saga,” in Proceedings of the 21st International Conference on Mining Software Repositories , 2024, pp. 177–181
2024
Closest in time.