Fetching the paper…
Reading the bibliography…
Enterprise penetration-testing is often limited by high operational costs and the scarcity of human expertise.
Heuristic evaluation of user interfaces. In Proceedings of the SIGCHI conference on Human factors in computing systems . 249–256
Jakob Nielsen and Rolf Molich. 1990 · 1990
Earlier work this paper cites.
Preliminary guidelines for empirical research in software engineering
Barbara A Kitchenham, Shari Lawrence Pfleeger, Lesley M Pickard, Peter W Jones, David C. Hoaglin, Khaled El Emam, and Jarrett Rosenberg. 2002 · 2002
Earlier work this paper cites.
Mastering Active directory for Windows server 2003
Robert R King. 2006 · 2003
Earlier work this paper cites.
Using thematic analysis in psychology
Virginia Braun and Victoria Clarke. 2006 · 2006
Earlier work this paper cites.
Constructing grounded theory: A practical guide through qualitative analysis
Kathy Charmaz. 2006 · 2006
Earlier work this paper cites.
Outside the closed world: On using machine learning for network intrusion detection. In 2010 IEEE symposium on security and privacy . IEEE, 305–316
Robin Sommer and Vern Paxson. 2010 · 2010
Earlier work this paper cites.
POMDPs make better hackers: Accounting for uncertainty in penetration testing. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 26. 1816–1824
Carlos Sarraute, Olivier Buffet, and Jörg Hoffmann. 2012 · 2012
Earlier work this paper cites.
Penetration testing== POMDP solving?
Carlos Sarraute, Olivier Buffet, and Jörg Hoffmann. 2013 · 2013
Earlier work this paper cites.
Sociological methods: A sourcebook
Norman K Denzin. 2017 · 2017
Earlier work this paper cites.
Active Directory and Related Aspects of Security. In 2018 21st Saudi Computer Society National Computer Conference (NCC) . 4474–4479
Afnan Binduf, Hanan Othman Alamoudi, Hanan Balahmar, Shatha Alshamrani, Haifa Al-Omar, and Naya Nagy. 2018 · 2018
Earlier work this paper cites.
Zero trust architecture
V Stafford. 2020 · 2020
Earlier work this paper cites.
The rise of ransomware: Forensic analysis for windows based ransomware attacks
Ilker Kara and Murat Aydos. 2022 · 2021
Earlier work this paper cites.
Caldera: A red-blue cyber operations automation platform
Ron Alford, Dean Lawrence, and Michael Kouremetis. 2022 · 2022
Earlier work this paper cites.
Sample sizes for saturation in qualitative research: A systematic review of empirical tests
Monique Hennink and Bonnie N Kaiser. 2022 · 2022
Earlier work this paper cites.
The Rise of Ransomware: A Review of Attacks, Detection Techniques, and Future Challenges. In 2022 International Conference on Business Analytics for Technology and Security (ICBATS) . 1–7
Samar Kamil, Huda Sheikh Abdullah Siti Norul, Ahmad Firdaus, and Opeyemi Lateef Usman. 2022 · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022 · 2022
Earlier work this paper cites.
Comparing attack models for it systems: Lockheed martin’s cyber kill chain, mitre att&ck framework and diamond model. In 2022 IEEE International Symposium on Systems Engineering (ISSE) . IEEE, 1–7
Nitin Naik, Paul Jenkins, Paul Grace, and Jingping Song. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Earlier work this paper cites.
Concepts and evaluation of saturation in qualitative research
Liping Yang, Lidong QI, and Bo Zhang. 2022 · 2022
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022 · 2022
Earlier work this paper cites.
Automatic chain of thought prompting in large language models
Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. 2022 · 2022
Earlier work this paper cites.
Penheal: a two-stage llm framework for automated pentesting and optimal remediation. In Proceedings of the Workshop on Autonomous Cybersecurity . 11–22
Junjie Huang and Quanyan Zhu. 2023 · 2023
Cited alongside, same era.
SoK: The MITRE ATT&CK Framework in Research and Practice
Shanto Roy, Emmanouil Panaousis, Cameron Noakes, Aron Laszka, Sakshyam Panda, and George Loukas. 2023 · 2023
Cited alongside, same era.
Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models
Lei Wang, Wanyu Xu, Yihuai Lan, Zhiqiang Hu, Yunshi Lan, Roy Ka-Wei Lee, and Ee-Peng Lim. 2023 · 2023
Cited alongside, same era.
PentestGPT: An LLM-empowered Automatic Penetration Testing Tool
Gelei Deng, Yi Liu, Víctor Mayoral-Vilches, Peng Liu, Yuekang Li, Yuan Xu, Tianwei Zhang, Yang Liu, Martin Pinzger, and Stefan Rass. 2024 · 2024
Cited alongside, same era.
VulnBot: Autonomous Penetration Testing for A Multi-Agent Collaborative Framework
He Kong, Die Hu, Jingguo Ge, Liangxiong Li, Tong Li, and Bingzhen Wu. 2025 · 2025
Closest in time.
Active Directory Holds the Keys to your Kingdom, but is it Secure?
Swetha Krishnamoorthi and Jarad Carleton. 2020 · 2025
Closest in time.
When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs
Xiaomin Li, Zhou Yu, Zhiwei Zhang, Xupeng Chen, Ziji Zhang, Yingying Zhuang, Narayanan Sadagopan, and Anurag Beniwal. 2025 · 2025
Closest in time.
LLM Cyber Evaluations Don’t Capture Real-World Risk
Kamilė Lukošiūtė and Adam Swanda. 2025 · 2025
Closest in time.
Global Ransomware Damage Costs Predicted To Exceed $275 Billion By 2031
Steve Morgan. 2025 · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Luca Gioacchini, Marco Mellia, Idilio Drago, Alexander Delsanto, Giuseppe Siracusano, and Roberto Bifulco. 2024 · 2024
Cited alongside, same era.
Llms as hackers: Autonomous linux privilege escalation attacks
Andreas Happe, Aaron Kaplan, and Juergen Cito. 2024 · 2024
Cited alongside, same era.
Fred Heiding, Simon Lermen, Andrew Kao, Bruce Schneier, and Arun Vishwanath. 2024a · 2024
Cited alongside, same era.
Towards automated penetration testing: Introducing llm benchmark, analysis, and improvements
Isamu Isozaki, Manil Shrestha, Rick Console, and Edward Kim. 2024 · 2024
Cited alongside, same era.
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al · 2024
Cited alongside, same era.
Evolution of endpoint detection and response (edr) in cyber security: A comprehensive review. In E3S Web of Conferences , Vol. 556. EDP Sciences, 01006
Harpreet Kaur, Dharani Sanjaiy SL, Tirtharaj Paul, Rohit Kumar Thakur, K Vijay Kumar Reddy, Jay Mahato, and Kaviti Naveen. 2024 · 2024
Cited alongside, same era.
HackSynth: LLM Agent and Evaluation Framework for Autonomous Penetration Testing
Lajos Muzsai, David Imolai, and András Lukács. 2024 · 2024
Cited alongside, same era.
ChainReactor: Automated Privilege Escalation Chain Discovery via AI Planning. In 33rd USENIX Security Symposium (USENIX Security 24) . USENIX Association, Philadelphia, PA, 5913–5929
Giulio De Pasquale, Ilya Grishchenko, Riccardo Iesari, Gabriel Pizarro, Lorenzo Cavallaro, Christopher Kruegel, and Giovanni Vigna. 2024 · 2024
Cited alongside, same era.
Closest in time.
RapidPen: Fully Automated IP-to-Shell Penetration Testing with LLM-based Agents
Sho Nakatani. 2025 · 2025
Closest in time.
How ransomware happens and how to stop it
part of the National Cyber Security Centre (NCSC) New Zealand’s CERT (Computer Emergency Response Team). 2023 · 2025
Closest in time.
Introducing OpenAI o1-preview
OpenAI. 2024a · 2025
Closest in time.
Learning to reason with LLMs
OpenAI. 2024b · 2025
Closest in time.
As some of you have noticed, avoid “boomer prompts” with o-series models. Instead, be simple and direct, with specific guidelines
OpenAI. 2025a · 2025
Closest in time.
Reasoning best practices
OpenAI. 2025b · 2025
Closest in time.
Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad
Ivo Petrov, Jasper Dekoninck, Lyuben Baltadzhiev, Maria Drencheva, Kristian Minchev, Mislav Balunović, Nikola Jovanović, and Martin Vechev. 2025 · 2025
Closest in time.
BoomerPrompts
Boomer Prompts. 2025 · 2025
Closest in time.
Evaluating Agent-based Program Repair at Google
Pat Rondon, Renyao Wei, José Cambronero, Jürgen Cito, Aaron Sun, Siddhant Sanyam, Michele Tufano, and Satish Chandra. 2025 · 2025
Closest in time.
Attackers Set Sights on Active Directory: Understanding Your Identity Exposure
Venu Shastri. 2022 · 2025
Closest in time.
Parshin Shojaee, Iman Mirzadeh, Keivan Alizadeh, Maxwell Horton, Samy Bengio, and Mehrdad Farajtabar. 2025 · 2025
Closest in time.
On the Feasibility of Using LLMs to Execute Multistage Network Attacks
Brian Singer, Keane Lucas, Lakshmi Adiga, Meghna Jain, Lujo Bauer, and Vyas Sekar. 2025 · 2025
Closest in time.
25 Years On, Active Directory Is Still a Prime Attack Target
Jai Vijayan. 2025 · 2025
Closest in time.
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al · 2025
Closest in time.
Getting pwn’d by AI: Penetration Testing with Large Language Models. In Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE ’23) . ACM, 2082–2086
Andreas Happe and Jürgen Cito. 2023a · 2086
Closest in time.