Fetching the paper…
Reading the bibliography…
Security operators use red teams to simulate real attackers and proactively find defense gaps.
Automated generation and analysis of attack graphs
Oleg Sheyner, Joshua Haines, Somesh Jha, Richard Lippmann, and Jeannette M Wing · 2002
Earlier work this paper cites.
Mulval: A logic-based network security analyzer
Xinming Ou, Sudhakar Govindavajhala, Andrew W Appel, et al · 2005
Earlier work this paper cites.
A scalable approach to attack graph generation
Xinming Ou, Wayne F Boyer, and Miles A McQueen · 2006
Earlier work this paper cites.
Technical report, Cisco, April 2008
Enterprise Campus 3.0 Architecture: Overview and Framework · 2008
Earlier work this paper cites.
Nmap network scanning: The official Nmap project guide to network discovery and security scanning
Gordon Fyodor Lyon · 2009
Earlier work this paper cites.
Before we knew it: an empirical study of zero-day attacks in the real world
Leyla Bilge and Tudor Dumitraş · 2012
Earlier work this paper cites.
A technique for network topology deception
Samuel T Trassare, Robert Beverly, and David Alderson · 2013
Earlier work this paper cites.
Intelligent, Automated Red Team Emulation
Andy Applebaum, Doug Miller, Blake Strom, Chris Korban, and Ross Wolf · 2016
Earlier work this paper cites.
SVED: Scanning, vulnerabilities, exploits and detection
Hannes Holm and Teodor Sommestad · 2016
Earlier work this paper cites.
CRASHOVERRIDE: Analysis of the threat to electric grid operations
Robert M Lee, MJ Assante, and T Conway · 2017
Earlier work this paper cites.
Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Andy K Zhang, Neil Perry, Riya Dulepet, Joey Ji, Justin W Lin, Eliot Jones, Celeste Menders, Gashon Hussein, Samantha Liu, Donovan Jasper, et al · 2017
Earlier work this paper cites.
The Equifax Data Breach
Majority Staff Report 115th Congress · 2018
Earlier work this paper cites.
DeepExploit: Fully Automatic Penetration Test Tool Using Reinforcement Learning
Isao Takaesu · 2018
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown · 2020
Earlier work this paper cites.
Harmer: Cyber-attacks automation and evaluation
Simon Yusuf Enoch, Zhibin Huang, Chun Yong Moon, Donghwan Lee, Myung Kil Ahn, and Dong Seong Kim · 2020
Earlier work this paper cites.
Automated penetration testing using deep reinforcement learning
Zhenguo Hu, Razvan Beuran, and Yutaka Tan · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Cyber attack and defense emulation agents
Jeong Do Yoo, Eunji Park, Gyungmin Lee, Myung Kil Ahn, Donghwa Kim, Seongyun Seo, and Huy Kang Kim · 2020
Earlier work this paper cites.
Offensive security: Towards proactive threat hunting via adversary emulation
Abdul Basit Ajmal, Munam Ali Shah, Carsten Maple, Muhammad Nabeel Asghar, and Saif Ul Islam · 2021
Earlier work this paper cites.
Examining the efficacy of decoy-based and psychological cyber deception
Kimberly J Ferguson-Walter, Maxine M Major, Chelsea K Johnson, and Daniel H Muhleman · 2021
Earlier work this paper cites.
Synchronizability of double-layer dumbbell networks
Juyi Li, Yangyang Luan, Xiaoqun Wu, and Jun-an Lu · 2021
Earlier work this paper cites.
Shining a light on darkside ransomware operations
Jordan Nuce, Jeremy Kennelly, Kimberly Goody, Andrew Moore, Alyssa Rahman, Matt Williams, Brendan McKeague, and Jared Wilson · 2021
Cited alongside, same era.
Equifax Data Breach Settlement
FTC · 2022
Cited alongside, same era.
Lore a red team emulation tool
Hannes Holm · 2022
Cited alongside, same era.
Colonial Pipeline hack explained: Everything you need to know
Sean Michael Kerner · 2022
Cited alongside, same era.
Introduction to ICS security fundamentals
Stephen Mathezer · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Cited alongside, same era.
PenHeal: A Two-Stage LLM Framework for Automated Pentesting and Optimal Remediation
Junjie Huang and Quanyan Zhu · 2024
Later among the works it cites.
Mirage: cyber deception against autonomous cyber attacks in emulation and simulation
Michael Kouremetis, Dean Lawrence, Ron Alford, Zoe Cheuvront, David Davila, Benjamin Geyer, Trevor Haigh, Ethan Michalak, Rachel Murphy, and Gianpaolo Russo · 2024
Later among the works it cites.
Chameleon: Plug-and-play compositional reasoning with large language models
Pan Lu, Baolin Peng, Hao Cheng, Michel Galley, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, and Jianfeng Gao · 2024
Later among the works it cites.
On sms phishing tactics and infrastructure
Aleksandr Nahapetyan, Sathvik Prasad, Kevin Childs, Adam Oest, Yeganeh Ladwig, Alexandros Kapravelos, and Bradley Reaves · 2024
Later among the works it cites.
OpenAI o1 System Card
OpenAI · 2024
Later among the works it cites.
From Chatbots to Phishbots?: Phishing Scam Generation in Commercial Large Language Models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Microsoft Security Copilot, 2023
Microsoft Corporation · 2023
Cited alongside, same era.
Semantic anomaly detection with large language models
Amine Elhafsi, Rohan Sinha, Christopher Agia, Edward Schmerling, Issa AD Nesnas, and Marco Pavone · 2023
Cited alongside, same era.
Pal: Program-aided language models
Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig · 2023
Cited alongside, same era.
Getting pwn’d by ai: Penetration testing with large language models
Andreas Happe and Jürgen Cito · 2023
Cited alongside, same era.
A Good Fishman Knows All the Angles: A Critical Evaluation of Google’s Phishing Page Classifier
Changqing Miao, Jianan Feng, Wei You, Wenchang Shi, Jianjun Huang, and Bin Liang · 2023
Cited alongside, same era.
Shedding light on inconsistencies in grid cybersecurity: Disconnects and recommendations
Brian Singer, Amritanshu Pandey, Shimiao Li, Lujo Bauer, Craig Miller, Lawrence Pileggi, and Vyas Sekar · 2023
Cited alongside, same era.
Sayak Saha Roy, Poojitha Thota, Krishna Vamsi Naragam, and Shirin Nilizadeh · 2024
Later among the works it cites.
Toolformer: Language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom · 2024
Later among the works it cites.
SoK: On the offensive potential of AI
Saskia Laura SchrÃķer, Giovanni Apruzzese, Soheil Human, Pavel Laskov, Hyrum S Anderson, Edward WN Bernroider, Aurore Fass, Ben Nassi, Vera Rimmer, Fabio Roli, et al · 2024
Later among the works it cites.
An empirical evaluation of llms for solving offensive security challenges
Minghao Shao, Boyuan Chen, Sofija Jancheska, Brendan Dolan-Gavitt, Siddharth Garg, Ramesh Karri, and Muhammad Shafique · 2024
Later among the works it cites.
Nyu ctf dataset: A scalable open-source benchmark dataset for evaluating llms in offensive security
Minghao Shao, Sofija Jancheska, Meet Udeshi, Brendan Dolan-Gavitt, Haoran Xi, Kimberly Milner, Boyuan Chen, Max Yin, Siddharth Garg, Prashanth Krishnamurthy, et al · 2024
Later among the works it cites.
Cybermetric: A benchmark dataset for evaluating large language models knowledge in cybersecurity
Norbert Tihanyi, Mohamed Amine Ferrag, Ridhi Jain, and Mer-ouane Debbah · 2024
Later among the works it cites.
US AISI and UK AISI Joint Pre-Deployment Test: Anthropic’s Claude 3.5 Sonnet
US AI Safety Institute and UK AI Safety Institute · 2024
Later among the works it cites.
Shengye Wan, Cyrus Nikolaidis, Daniel Song, David Molnar, James Crnkovich, Jayson Grace, Manish Bhatt, Sahana Chennabasappa, Spencer Whitman, Stephanie Ding, et al · 2024
Later among the works it cites.
Autoattacker: A large language model guided system to implement automatic cyber-attacks
Jiacen Xu, Jack W Stokes, Geoff McDonald, Xuesong Bai, David Marshall, Siyue Wang, Adith Swaminathan, and Zhou Li · 2024
Later among the works it cites.
https://agi.safe.ai/ , accessed: Apr 23, 2025
Humanity’s Last Exam · 2025
Closest in time.
Pentest++: Elevating ethical hacking with ai and automation
Haitham S Al-Sinani and Chris J Mitchell · 2025
Closest in time.
Vulnbot: Autonomous penetration testing for a multi-agent collaborative framework
He Kong, Die Hu, Jingguo Ge, Liangxiong Li, Tong Li, and Bingzhen Wu · 2025
Closest in time.
Occult: Evaluating large language models for offensive cyber operation capabilities
Michael Kouremetis, Marissa Dotter, Alex Byrne, Dan Martin, Ethan Michalak, Gianpaolo Russo, Michael Threet, and Guido Zarrella · 2025
Closest in time.
A framework for evaluating emerging cyberattack capabilities of ai
Mikel Rodriguez, Raluca Ada Popa, Four Flynn, Lihao Liang, Allan Dafoe, and Anna Wang · 2025
Closest in time.
Hashcat: Advanced Password Recovery
Jens Steube · 2025
Closest in time.