Fetching the paper…
Reading the bibliography…
We saturate a high-school-level hacking benchmark with plain LLM agent design.
“FACT SHEET: President Biden Issues Executive Order on Safe, Secure, and Trustworthy Artificial Intelligence”
The White House · 2023
Earlier work this paper cites.
“The Bletchley Declaration by Countries Attending the AI Safety Summit, 1-2 November 2023”
UK Government · 2023
Earlier work this paper cites.
“InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback”
John Yang, Akshara Prabhakar, Karthik Narasimhan and Shunyu Yao · 2023
Earlier work this paper cites.
“Language Agents as Hackers: Evaluating Cybersecurity Skills with Capture the Flag”
John Yang et al · 2023
Earlier work this paper cites.
“Tree of Thoughts: Deliberate Problem Solving with Large Language Models”
Shunyu Yao et al · 2023
Earlier work this paper cites.
“ReAct: Synergizing Reasoning and Acting in Language Models”
Shunyu Yao et al · 2023
Cited alongside, same era.
“EnIGMA: Enhanced Interactive Generative Model Agent for CTF Challenges”
Talor Abramovich et al · 2024
Cited alongside, same era.
“Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities”
Andrey Anurin et al · 2024
Cited alongside, same era.
“CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models”
Manish Bhatt et al · 2024
Cited alongside, same era.
“GPT-4o System Card”
OpenAI et al · 2024
Cited alongside, same era.
“SITUATIONAL AWARENESS: The Decade Ahead”, https://situational-awareness.ai/
Leopold Aschenbrenner
“Evaluating Frontier Models for Dangerous Capabilities”
Mary Phuong et al · 2024
Closest in time.
“Project Naptime: Evaluating Offensive Security Capabilities of Large Language Models”, https://googleprojectzero.blogspot.com/2024/06/project-naptime.html
Project Zero · 2024
Closest in time.
Minghao Shao et al · 2024
Closest in time.
“Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models”
Andy. Zhang et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited in the paper.
“An Update on Disrupting Deceptive Uses of AI”, https://openai.com/global-affairs/an-update-on-disrupting-deceptive-uses-of-ai/
OpenAI
Cited in the paper.
“Preparedness”, https://openai.com/safety/preparedness
OpenAI
Cited in the paper.