Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have emerged as a powerful approach for driving offensive penetration-testing tooling.
Attack trees
Bruce Schneier · 1999
Earlier work this paper cites.
Real world research, 2002
Collin Robson · 2002
Earlier work this paper cites.
Goodhart’s law: its origins, meaning and implications for monetary policy
K Alec Chrystal, Paul D Mizen, and PD Mizen · 2003
Earlier work this paper cites.
Using thematic analysis in psychology
Virginia Braun and Victoria Clarke · 2006
Earlier work this paper cites.
Peer debriefing
Valerie J Janesick · 2007
Earlier work this paper cites.
Sp 800-115. technical guide to information security testing and assessment
Karen A. Scarfone, Murugiah P. Souppaya, Amanda Cody, and Angela D. Orebaugh · 2008
Earlier work this paper cites.
Outside the closed world: On using machine learning for network intrusion detection
Robin Sommer and Vern Paxson · 2010
Earlier work this paper cites.
Extracting attack narratives from traffic datasets
Jose David Mireles, Jin-Hee Cho, and Shouhuai Xu · 2016
Earlier work this paper cites.
The red queen problem, innovation in the dod and intelligence community
Steve Blank · 2017
Earlier work this paper cites.
Thematic analysis: Striving to meet the trustworthiness criteria
Lorelli S Nowell, Jill M Norris, Deborah E White, and Nancy J Moules · 2017
Earlier work this paper cites.
A review of attack graph and attack tree visual syntax in cyber security
Harjinder Singh Lallie, Kurt Debattista, and Jay Bal · 2019
Earlier work this paper cites.
Standardised penetration testing? examining the usefulness of current penetration testing methodologies
Niek Jan van den Hout · 2019
Earlier work this paper cites.
An analysis and evaluation of open source capture the flag platforms as cybersecurity e-learning tools
Stylianos Karagiannis, Elpidoforos Maragkos-Belmpas, and Emmanouil Magkos · 2020
Earlier work this paper cites.
A capture the flag (ctf) platform and exercises for an intro to computer security class
Zack Kaplan, Ning Zhang, and Stephen V Cole · 2022
Cited alongside, same era.
Comparing attack models for it systems: Lockheed martin’s cyber kill chain, mitre att&ck framework and diamond model
Nitin Naik, Paul Jenkins, Paul Grace, and Jingping Song · 2022
Cited alongside, same era.
Purple llama cyberseceval: A secure coding benchmark for language models
Manish Bhatt, Sahana Chennabasappa, Cyrus Nikolaidis, Shengye Wan, Ivan Evtimov, Dominik Gabi, Daniel Song, Faizan Ahmad, Cornelius Aschermann, Lorenzo Fontana, et al · 2023
Cited alongside, same era.
Penheal: a two-stage llm framework for automated pentesting and optimal remediation
Junjie Huang and Quanyan Zhu · 2023
Cited alongside, same era.
Cyberseceval 2: A wide-ranging cybersecurity evaluation suite for large language models
Generative ai in cyber security of cyber physical systems: Benefits and threats
Harindra S. Mavikumbure, Victor Cobilean, Chathurika S. Wickramasinghe, Devin Drake, and Milos Manic · 2024
Later among the works it cites.
Large language models in cybersecurity: State-of-the-art, 2024
Farzad Nourmohammadzadeh Motlagh, Mehrdad Hajizadeh, Mehryar Majd, Pejman Najafi, Feng Cheng, and Christoph Meinel · 2024
Later among the works it cites.
Hacksynth: Llm agent and evaluation framework for autonomous penetration testing, 2024
Lajos Muzsai, David Imolai, and András Lukács · 2024
Later among the works it cites.
Shengye Wan, Cyrus Nikolaidis, Daniel Song, David Molnar, James Crnkovich, Jayson Grace, Manish Bhatt, Sahana Chennabasappa, Spencer Whitman, Stephanie Ding, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Manish Bhatt, Sahana Chennabasappa, Yue Li, Cyrus Nikolaidis, Daniel Song, Shengye Wan, Faizan Ahmad, Cornelius Aschermann, Yaohui Chen, Dhaval Kapil, et al · 2024
Cited alongside, same era.
Llms for intelligent software testing: A comparative study
Mohamed Boukhlif, Nassim Kharmoum, and Mohamed Hanine · 2024
Cited alongside, same era.
P e n t e s t G P T PentestGPT : Evaluating and harnessing large language models for automated penetration testing
Gelei Deng, Yi Liu, Víctor Mayoral-Vilches, Peng Liu, Yuekang Li, Yuan Xu, Tianwei Zhang, Yang Liu, Martin Pinzger, and Stefan Rass · 2024
Cited alongside, same era.
Large language models in information security research: A january 2024 survey
Rohit Dube · 2024
Cited alongside, same era.
Autopenbench: Benchmarking generative agents for penetration testing, 2024
Luca Gioacchini, Marco Mellia, Idilio Drago, Alexander Delsanto, Giuseppe Siracusano, and Roberto Bifulco · 2024
Cited alongside, same era.
Llms as hackers: Autonomous linux privilege escalation attacks
Andreas Happe, Aaron Kaplan, and Juergen Cito · 2024
Cited alongside, same era.
Mohammed Hassanin and Nour Moustafa · 2024
Cited alongside, same era.
Towards automated penetration testing: Introducing llm benchmark, analysis, and improvements
Isamu Isozaki, Manil Shrestha, Rick Console, and Edward Kim · 2024
Cited alongside, same era.
Benlong Wu, Guoqiang Chen, Kejiang Chen, Xiuwei Shang, Jiapeng Han, Yanru He, Weiming Zhang, and Nenghai Yu · 2024
Later among the works it cites.
A survey on large language model (llm) security and privacy: The good, the bad, and the ugly
Yifan Yao, Jinhao Duan, Kaidi Xu, Yuanfang Cai, Zhibo Sun, and Yue Zhang · 2024
Later among the works it cites.
Review of generative ai methods in cybersecurity, 2024
Yagmur Yigit, William J Buchanan, Madjid G Tehrani, and Leandros Maglaras · 2024
Later among the works it cites.
Andreas Happe and Jürgen Cito · 2025
Closest in time.
Vulnbot: Autonomous penetration testing for a multi-agent collaborative framework
He Kong, Die Hu, Jingguo Ge, Liangxiong Li, Tong Li, and Bingzhen Wu · 2025
Closest in time.
Rapidpen: Fully automated ip-to-shell penetration testing with llm-based agents, 2025
Sho Nakatani · 2025
Closest in time.
On the feasibility of using llms to execute multistage network attacks
Brian Singer, Keane Lucas, Lakshmi Adiga, Meghna Jain, Lujo Bauer, and Vyas Sekar · 2025
Closest in time.
Getting pwn’d by ai: Penetration testing with large language models
Andreas Happe and Jürgen Cito · 2086
Closest in time.