Fetching the paper…
Reading the bibliography…
LLM agents have become increasingly sophisticated, especially in the realm of cybersecurity.
A privacy-preserving defense mechanism against request forgery attacks
Ben SY Fung and Patrick PC Lee. 2011 · 2011
Earlier work this paper cites.
Metasploit: the penetration tester’s guide
David Kennedy, Jim O’gorman, Devon Kearns, and Mati Aharoni. 2011 · 2011
Earlier work this paper cites.
Before we knew it: an empirical study of zero-day attacks in the real world
Leyla Bilge and Tudor Dumitraş. 2012 · 2012
Earlier work this paper cites.
Enhancing intelligent agents with episodic memory
Andrew M Nuxoll and John E Laird. 2012 · 2012
Earlier work this paper cites.
Owasp zed attack proxy
Simon Bennetts. 2013 · 2013
Earlier work this paper cites.
Web vulnerability analysis and implementation
Eko Budi Setiawan and Angga Setiyadi. 2018 · 2018
Earlier work this paper cites.
Machine learning in cybersecurity: A review
Anand Handa, Ashu Sharma, and Sandeep K Shukla. 2019 · 2019
Earlier work this paper cites.
Critical microsoft exchange flaw: What is cve-2021-26855?
Edward Kost. 2023 · 2021
Earlier work this paper cites.
Will ai make cyber swords or shields?
Andrew Lohn and Krystal Jackson. 2022 · 2022
Earlier work this paper cites.
Talm: Tool augmented language models
Aaron Parisi, Yao Zhao, and Noah Fiedel. 2022 · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022 · 2022
Earlier work this paper cites.
ReAct: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022 · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Earlier work this paper cites.
Autoagents: A framework for automatic agent generation
Guangyao Chen, Siwei Dong, Yu Shu, Ge Zhang, Jaward Sesay, Börje F Karlsson, Jie Fu, and Yemin Shi. 2023 · 2023
Cited alongside, same era.
Exploiting programmatic behavior of llms: Dual-use through standard security attacks
Daniel Kang, Xuechen Li, Ion Stoica, Carlos Guestrin, Matei Zaharia, and Tatsunori Hashimoto. 2023 · 2023
Cited alongside, same era.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. 2023 · 2023
Cited alongside, same era.
Llm-powered autonomous agents
Lilian Weng. 2023 · 2023
Cited alongside, same era.
Shadow alignment: The ease of subverting safely-aligned language models
Securing the cloud
Microsoft. 2024 · 2024
Closest in time.
Vulnerability disclosure cheat sheet
OWASP. 2024 · 2024
Closest in time.
Google i/o 2024: everything announced
Emma Roth and Wes Davis. 2024 · 2024
Closest in time.
Reflexion: Language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2024 · 2024
Closest in time.
sqlmap: Automatic sql injection and database takeover tool
Project sqlmap. 2024 · 2024
Closest in time.
Ai safety institute approach to evaluations
AISI UK. 2024 · 2024
Closest in time.
Holistic safety and responsibility evaluations of advanced ai models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xianjun Yang, Xiao Wang, Qi Zhang, Linda Petzold, William Yang Wang, Xun Zhao, and Dahua Lin. 2023 · 2023
Cited alongside, same era.
Benchmarking and defending against indirect prompt injection attacks on large language models
Jingwei Yi, Yueqi Xie, Bin Zhu, Keegan Hines, Emre Kiciman, Guangzhong Sun, Xing Xie, and Fangzhao Wu. 2023 · 2023
Cited alongside, same era.
Removing rlhf protections in gpt-4 via fine-tuning
Qiusi Zhan, Richard Fang, Rohan Bindu, Akul Gupta, Tatsunori Hashimoto, and Daniel Kang. 2023 · 2023
Cited alongside, same era.
Building cooperative embodied agents modularly with large language models
Hongxin Zhang, Weihua Du, Jiaming Shan, Qinhong Zhou, Yilun Du, Joshua B Tenenbaum, Tianmin Shu, and Chuang Gan. 2023 · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson. 2023 · 2023
Cited alongside, same era.
A new initiative for developing third-party model evaluations
Anthropic. 2024 · 2024
Cited alongside, same era.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024 · 2024
Cited alongside, same era.
Webvoyager: Building an end-to-end web agent with large multimodal models
Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu. 2024 · 2024
Cited alongside, same era.
Laura Weidinger, Joslyn Barnhart, Jenny Brennan, Christina Butterfield, Susie Young, Will Hawkins, Lisa Anne Hendricks, Ramona Comanescu, Oscar Chang, Mikel Rodriguez, et al. 2024 · 2024
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2024 · 2024
Closest in time.
Injecagent: Benchmarking indirect prompt injections in tool-integrated large language model agents
Qiusi Zhan, Zhixiang Liang, Zifan Ying, and Daniel Kang. 2024 · 2024
Closest in time.
Cybench: A framework for evaluating cybersecurity capabilities and risks of language models
Andy K Zhang, Neil Perry, Riya Dulepet, Joey Ji, Justin W Lin, Eliot Jones, Celeste Menders, Gashon Hussein, Samantha Liu, Donovan Jasper, et al. 2024 · 2024
Closest in time.
Technical blog: Strengthening ai agent hijacking evaluations
AISI US. 2025 · 2025
Closest in time.
Getting pwn’d by ai: Penetration testing with large language models
Andreas Happe and Jürgen Cito. 2023 · 2086
Closest in time.