Fetching the paper…
Reading the bibliography…
Recent advances in AI agents capable of solving complex, everyday tasks, from scheduling to customer service, have enabled deployment in real-world settings, but their possibilities for unsafe behavior demands rigorous evaluation.
“Certifying llm safety against adversarial prompting”
Aounon Kumar et al · 2023
Earlier work this paper cites.
“Self-guard: Empower the llm to safeguard itself”
Zezhong Wang et al · 2023
Earlier work this paper cites.
“WebArena: A Realistic Web Environment for Building Autonomous Agents”
Shuyan Zhou et al · 2023
Earlier work this paper cites.
“DeepSeek-V3 Technical Report” Accessed: 2025-05-04, https://arxiv.org/abs/2412.19437 , 2024
DeepSeek-AI et al · 2024
Earlier work this paper cites.
“RedCode: Risky Code Execution and Generation Benchmark for Code Agents”, 2024
Chengquan Guo et al · 2024
Earlier work this paper cites.
“WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models”, 2024
Liwei Jiang et al · 2024
Earlier work this paper cites.
“Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents”, 2024
Priyanshu Kumar et al · 2024
Earlier work this paper cites.
“State of AI Agents 2024 Report” Survey of over 1,300 professionals on AI agent adoption across industries, https://www.langchain.com/stateofaiagents , 2024
LangChain · 2024
Earlier work this paper cites.
“ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents”, 2024
Ido Levy et al · 2024
Earlier work this paper cites.
“SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models”, 2024
Lijun Li et al · 2024
Earlier work this paper cites.
“A Trembling House of Cards? Mapping Adversarial Attacks against Language Agents”, 2024
Lingbo Mo et al · 2024
Earlier work this paper cites.
“SG-Bench: Evaluating LLM Safety Generalization Across Diverse Tasks and Prompt Types”, 2024
Yutao Mou, Shikun Zhang and Wei Ye · 2024
Earlier work this paper cites.
OpenAI et al · 2024
Earlier work this paper cites.
“Identifying the Risks of LM Agents with an LM-Emulated Sandbox”, 2024
Yangjun Ruan et al · 2024
Earlier work this paper cites.
“Prioritizing Safeguarding Over Autonomy: Risks of LLM Agents for Science”, 2024
Xiangru Tang et al · 2024
Earlier work this paper cites.
Simone Tedeschi et al · 2024
Earlier work this paper cites.
“AdvWeb: Controllable Black-box Attacks on VLM-powered Web Agents”, 2024
Chejian Xu et al · 2024
Cited alongside, same era.
“TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks”, 2024
Frank. Xu et al · 2024
Cited alongside, same era.
“SafeBench: A Safety Evaluation Framework for Multimodal Large Language Models”, 2024
Zonghao Ying et al · 2024
Cited alongside, same era.
“R-Judge: Benchmarking Safety Risk Awareness for LLM Agents”, 2024
Tongxin Yuan et al · 2024
Cited alongside, same era.
“AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies”, 2024
“h4rm3l: A language for Composable Jailbreak Attack Synthesis”, 2025
Moussa Doumbouya et al · 2025
Closest in time.
Philipp Guldimann et al · 2025
Closest in time.
“DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning”, 2025
Daya Guo et al · 2025
Closest in time.
“Commercial LLM Agents Are Already Vulnerable to Simple Yet Dangerous Attacks”, 2025
Ang Li et al · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yi Zeng et al · 2024
Cited alongside, same era.
“Attacking Vision-Language Computer Agents via Pop-ups”, 2024
Yanzhe Zhang, Tao Yu and Diyi Yang · 2024
Cited alongside, same era.
“Agent-SafetyBench: Evaluating the Safety of LLM Agents”, 2024
Zhexin Zhang et al · 2024
Cited alongside, same era.
“ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain”, 2024
Haochen Zhao et al · 2024
Cited alongside, same era.
“HAICOSYSTEM: An Ecosystem for Sandboxing Safety Risks in Human-AI Interactions”, 2024
Xuhui Zhou et al · 2024
Cited alongside, same era.
“SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents”, 2024
Xuhui Zhou et al · 2024
Cited alongside, same era.
“SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents”
Xuhui Zhou* et al · 2024
Cited alongside, same era.
“Firewalls to Secure Dynamic LLM Agentic Networks”, 2025
Sahar Abdelnabi et al · 2025
Cited alongside, same era.
Alexander Meinke et al · 2025
Closest in time.
“GPT-4.1” Large language model. Released April 14, 2025, https://openai.com/index/gpt-4-1/ , 2025
OpenAI · 2025
Closest in time.
“Agentic Large Language Models, a survey”, 2025
Aske Plaat et al · 2025
Closest in time.
“AgentBreeder: Mitigating the AI Safety Impact of Multi-Agent Scaffolds via Self-Improvement”, 2025
J Rosser and Jakob Foerster · 2025
Closest in time.
Paul Röttger, Fabio Pernisi, Bertie Vidgen and Dirk Hovy · 2025
Closest in time.
“Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration”, 2025
Yijia Shao et al · 2025
Closest in time.
“PrivacyLens: Evaluating Privacy Norm Awareness of Language Models in Action”, 2025
Yijia Shao et al · 2025
Closest in time.
“SafeArena: Evaluating the Safety of Autonomous Web Agents”, 2025
Ada Tur et al · 2025
Closest in time.
“OpenHands: An Open Platform for AI Software Developers as Generalist Agents”, 2025
Xingyao Wang et al · 2025
Closest in time.
“Dissecting Adversarial Robustness of Multimodal LM Agents”, 2025
Chen Wu et al · 2025
Closest in time.
“SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents”, 2025
Sheng Yin et al · 2025
Closest in time.
“OpenAI o3-mini System Card” Please cite this work as “OpenAI (2025)”, https://cdn.openai.com/o3-mini-system-card-feb10.pdf , 2025
Brian Zhang et al · 2025
Closest in time.