Fetching the paper…
Reading the bibliography…
Despite the critical threat posed by software security vulnerabilities, reports are often incomplete, lacking the proof-of-vulnerability (PoV) tests needed to validate fixes and prevent regressions.
Automatic patch-based exploit generation is possible: Techniques and implications. In 2008 IEEE Symposium on Security and Privacy (sp 2008) . IEEE, 143–157
David Brumley, Pongsin Poosankam, Dawn Song, and Jiang Zheng. 2008 · 2008
Earlier work this paper cites.
Craxweb: Automatic web application testing and attack generation. In 2013 IEEE 7th International Conference on Software Security and Reliability . IEEE, 208–217
Shih-Kun Huang, Han-Lin Lu, Wai-Meng Leong, and Huan Liu. 2013 · 2013
Earlier work this paper cites.
Automatic exploit generation
Thanassis Avgerinos, Sang Kil Cha, Alexandre Rebert, Edward J Schwartz, Maverick Woo, and David Brumley. 2014 · 2014
Earlier work this paper cites.
Automatic Generation of { \{ Data-Oriented } \} Exploits. In 24th USENIX Security Symposium (USENIX Security 15) . 177–192
Hong Hu, Zheng Leong Chua, Sendroiu Adrian, Prateek Saxena, and Zhenkai Liang. 2015 · 2015
Earlier work this paper cites.
Chainsaw: Chained automated workflow-based exploit generation. In Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security . 641–652
Abeer Alhuzali, Birhanu Eshete, Rigel Gjomemo, and VN Venkatakrishnan. 2016 · 2016
Earlier work this paper cites.
{ \{ OSS-Fuzz } \} -Google’s continuous fuzzing service for open source software
Kostya Serebryany. 2017 · 2017
Earlier work this paper cites.
Semfuzz: Semantics-based automatic generation of proof-of-concept exploits. In Proceedings of the 2017 ACM SIGSAC conference on computer and communications security . 2139–2154
Wei You, Peiyuan Zong, Kai Chen, XiaoFeng Wang, Xiaojing Liao, Pan Bian, and Bin Liang. 2017 · 2017
Earlier work this paper cites.
Understanding the reproducibility of crowd-reported security vulnerabilities. In 27th USENIX Security Symposium (USENIX Security 18) . 919–936
Dongliang Mu, Alejandro Cuevas, Limin Yang, Hang Hu, Xinyu Xing, Bing Mao, and Gang Wang. 2018 · 2018
Earlier work this paper cites.
AC/C++ code vulnerability dataset with code changes and CVE summaries. In Proceedings of the 17th international conference on mining software repositories . 508–512
Jiahao Fan, Yi Li, Shaohua Wang, and Tien N Nguyen. 2020 · 2020
Earlier work this paper cites.
CrossVul: a cross-language vulnerability dataset with commit data. In Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering . 1565–1569
Georgios Nikitopoulos, Konstantina Dritsa, Panos Louridas, and Dimitris Mitropoulos. 2021 · 2021
Earlier work this paper cites.
Vul4J: a dataset of reproducible Java vulnerabilities geared towards the study of program repair techniques. In Proceedings of the 19th International Conference on Mining Software Repositories . 464–468
Quang-Cuong Bui, Riccardo Scandariato, and Nicolás E Díaz Ferreyra. 2022 · 2022
Earlier work this paper cites.
SecBench. js: An executable security benchmark suite for server-side JavaScript. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE) . IEEE, 1059–1070
Masudul Hasan Masud Bhuiyan, Adithya Srinivas Parthasarathy, Nikos Vasilakis, Michael Pradel, and Cristian-Alexandru Staicu. 2023 · 2023
Cited alongside, same era.
A survey on automated software vulnerability detection using machine learning and deep learning
Nima Shiri Harzevili, Alvine Boaye Belle, Junjie Wang, Song Wang, Zhen Ming, Nachiappan Nagappan, et al · 2023
Cited alongside, same era.
Large language models for code: Security hardening and adversarial testing. In Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security . 1865–1879
Jingxuan He and Martin Vechev. 2023 · 2023
Cited alongside, same era.
Swe-bench: Can language models resolve real-world github issues?
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. 2023 · 2023
Cited alongside, same era.
Swe-agent: Agent-computer interfaces enable automated software engineering
John Yang, Carlos E Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. 2024 · 2024
Later among the works it cites.
Autocoderover: Autonomous program improvement. In Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis . 1592–1604
Yuntong Zhang, Haifeng Ruan, Zhiyu Fan, and Abhik Roychoudhury. 2024 · 2024
Later among the works it cites.
Teams of llm agents can exploit zero-day vulnerabilities
Yuxuan Zhu, Antony Kellermann, Akul Gupta, Philip Li, Richard Fang, Rohan Bindu, and Daniel Kang. 2024 · 2024
Later among the works it cites.
Otter: Generating Tests from Issues to Validate SWE Patches
Toufique Ahmed, Jatin Ganhotra, Rangeet Pan, Avraham Shinnar, Saurabh Sinha, and Martin Hirzel. 2025 · 2025
Closest in time.
Agentic Bug Reproduction for Effective Automated Program Repair at Google
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved?
Toufique Ahmed, Martin Hirzel, Rangeet Pan, Avraham Shinnar, and Saurabh Sinha. 2024 · 2024
Cited alongside, same era.
Vulnerability detection with code language models: How far are we?
Yangruibo Ding, Yanjun Fu, Omniyyah Ibrahim, Chawin Sitawarin, Xinyun Chen, Basel Alomair, David Wagner, Baishakhi Ray, and Yizheng Chen. 2024 · 2024
Cited alongside, same era.
Llm agents can autonomously exploit one-day vulnerabilities
Richard Fang, Rohan Bindu, Akul Gupta, and Daniel Kang. 2024 · 2024
Cited alongside, same era.
Evaluating diverse large language models for automatic and general bug reproduction
Sungmin Kang, Juyeon Yoon, Nargiz Askarbekkyzy, and Shin Yoo. 2024 · 2024
Cited alongside, same era.
ARVO: Atlas of Reproducible Vulnerabilities for Open Source Software
Xiang Mei, Pulkit Singh Singaria, Jordi Del Castillo, Haoran Xi, Tiffany Bao, Ruoyu Wang, Yan Shoshitaishvili, Adam Doupé, Hammond Pearce, Brendan Dolan-Gavitt, et al · 2024
Cited alongside, same era.
OpenHands CodeAct 2.1: An Open, State-of-the-Art Software Development Agent
Graham Neubig and Xingyao Wang. 2024 · 2024
Cited alongside, same era.
Introducing OpenDevin CodeAct 1.0, a new State-of-the-art in Coding Agents
Xingyao Wang, Bowen Li, and Graham Neubig. 2024b · 2024
Cited alongside, same era.
https://codeql.github.com/
[n. d.]
Cited in the paper.
Runxiang Cheng, Michele Tufano, Jürgen Cito, José Cambronero, Pat Rondon, Renyao Wei, Aaron Sun, and Satish Chandra. 2025 · 2025
Closest in time.
IRIS: LLM-assisted static analysis for detecting security vulnerabilities. In The Thirteenth International Conference on Learning Representations
Ziyang Li, Saikat Dutta, and Mayur Naik. 2025 · 2025
Closest in time.
Common Weakness Enumeration
MITRE Corporation. 2025 · 2025
Closest in time.
National Vulnerability Database
National Institute of Standards and Technology. 2025 · 2025
Closest in time.
PoCGen: Generating Proof-of-Concept Exploits for Vulnerabilities in Npm Packages
Deniz Simsek, Aryaz Eghbali, and Michael Pradel. 2025 · 2025
Closest in time.
OpenHands: An Open Platform for AI Software Developers as Generalist Agents. In The Thirteenth International Conference on Learning Representations
Xingyao Wang, Boxuan Li, Yufan Song, Frank F. Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H. Tran, Fuqiang Li, Ren Ma, Mingzhang Zheng, Bill Qian, Yanjun Shao, Niklas Muennighoff, Yizhe Zhang, Binyuan Hui, Junyang Lin, Robert Brennan, Hao Peng, Heng Ji, and Graham Neubig. 2025 · 2025
Closest in time.
CVE-Bench: A Benchmark for AI Agents’ Ability to Exploit Real-World Web Application Vulnerabilities
Yuxuan Zhu, Antony Kellermann, Dylan Bowman, Philip Li, Akul Gupta, Adarsh Danda, Richard Fang, Conner Jensen, Eric Ihli, Jason Benn, et al · 2025
Closest in time.