Fetching the paper…
Reading the bibliography…
AI agents interact with external environments through tool calls, exposing them to attacks like indirect prompt injection that can trigger unauthorized actions.
Binder, a logic-based security language. In Proceedings 2002 IEEE Symposium on Security and Privacy . IEEE, 105–113
John DeTreville. 2002 · 2002
Earlier work this paper cites.
Z3: An efficient SMT solver. In TACAS
Leonardo De Moura and Nikolaj Bjørner. 2008 · 2008
Earlier work this paper cites.
Sapper: A language for hardware-level security policy enforcement. In Proceedings of the 19th international conference on Architectural support for programming languages and operating systems . 97–112
Xun Li, Vineeth Kashyap, Jason K Oberg, Mohit Tiwari, Vasanth Ram Rajarathinam, Ryan Kastner, Timothy Sherwood, Ben Hardekopf, and Frederic T Chong. 2014 · 2014
Earlier work this paper cites.
Deberta: Decoding-enhanced bert with disentangled attention. In ICLR
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021 · 2021
Earlier work this paper cites.
Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security . 79–90
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. 2023 · 2023
Earlier work this paper cites.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, et al · 2023
Earlier work this paper cites.
Prompt Injection attack against LLM-integrated Applications
Yi Liu, Gelei Deng, Yuekang Li, Kailong Wang, Zihao Wang, Xiaofeng Wang, Tianwei Zhang, Yepang Liu, Haoyu Wang, Yan Zheng, et al · 2023
Earlier work this paper cites.
Toolllm: Facilitating large language models to master 16000+ real-world apis
Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, et al · 2023
Earlier work this paper cites.
Toolformer: Language models can teach themselves to use tools. In NeurIPS
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023 · 2023
Earlier work this paper cites.
Reflexion: Language agents with verbal reinforcement learning. In NeurIPS
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023 · 2023
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models. In ICLR
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023 · 2023
Earlier work this paper cites.
Airgapagent: Protecting privacy-conscious conversational agents. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security . 3868–3882
Eugene Bagdasarian, Ren Yi, Sahra Ghalebikesabi, Peter Kairouz, Marco Gruteser, Sewoong Oh, Borja Balle, and Daniel Ramage. 2024 · 2024
Earlier work this paper cites.
Cedar: A new language for expressive, fast, safe, and analyzable authorization
Joseph W Cutler, Craig Disselkoen, Aaron Eline, Shaobo He, Kyle Headley, Michael Hicks, Kesha Hietala, Eleftherios Ioannidis, John Kastner, Anwar Mamat, et al · 2024
Earlier work this paper cites.
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track
Edoardo Debenedetti, Jie Zhang, Mislav Balunovic, Luca Beurer-Kellner, Marc Fischer, and Florian Tramèr. 2024 · 2024
Earlier work this paper cites.
The emerged security and privacy of llm agent: A survey with case studies
Feng He, Tianqing Zhu, Dayong Ye, Bo Liu, Wanlei Zhou, and Philip S Yu. 2024 · 2024
Earlier work this paper cites.
Defending against indirect prompt injection attacks with spotlighting
Keegan Hines, Gary Lopez, Matthew Hall, Federico Zarfati, Yonatan Zunger, and Emre Kiciman. 2024 · 2024
Earlier work this paper cites.
Rongchang Li, Minjie Chen, Chang Hu, Han Chen, Wenpeng Xing, and Meng Han. 2024a · 2024
Earlier work this paper cites.
Personal llm agents: Insights and survey about the capability, efficiency and security
Yuanchun Li, Hao Wen, Weijun Wang, Xiangyu Li, Yizhen Yuan, Guohong Liu, Jiacheng Liu, Wenxing Xu, Xiang Wang, Yi Sun, et al · 2024
Earlier work this paper cites.
Fine-Tuned DeBERTa-v3-base for Prompt Injection Detection
ProtectAI.com. 2024 · 2024
Cited alongside, same era.
Ehragent: Code empowers large language models for few-shot complex tabular reasoning on electronic health records. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing . 22315–22339
Wenqi Shi, Ran Xu, Yuchen Zhuang, Yue Yu, Jieyu Zhang, Hang Wu, Yuanda Zhu, Joyce Ho, Carl Yang, and May Dongmei Wang. 2024 · 2024
Cited alongside, same era.
The instruction hierarchy: Training llms to prioritize privileged instructions
Eric Wallace, Kai Xiao, Reimar Leike, Lilian Weng, Johannes Heidecke, and Alex Beutel. 2024 · 2024
Cited alongside, same era.
A survey on large language model based autonomous agents
Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al · 2024
Cited alongside, same era.
Dissecting Adversarial Robustness of Multimodal LM Agents. In NeurIPS 2024 Workshop on Open-World Agents
Instruction Defense
Learn Prompting. 2024a · 2025
Closest in time.
Random Sequence Enclosure
Learn Prompting. 2024b · 2025
Closest in time.
Sandwich Defense
Learn Prompting. 2024c · 2025
Closest in time.
Drift: Dynamic rule-based defense with injection isolation for securing llm agents
Hao Li, Xiaogeng Liu, Hung-Chun Chiu, Dianqi Li, Ning Zhang, and Chaowei Xiao. 2025 · 2025
Closest in time.
Eia: Environmental injection attack on generalist web agents for privacy leakage
Zeyi Liao, Lingbo Mo, Chejian Xu, Mintong Kang, Jiawei Zhang, Chaowei Xiao, Yuan Tian, Bo Li, and Huan Sun. 2025 · 2025
Closest in time.
DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks
Yupei Liu, Yuqi Jia, Jinyuan Jia, Dawn Song, and Neil Zhenqiang Gong. 2025 · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen Henry Wu, Rishi Rajesh Shah, Jing Yu Koh, Russ Salakhutdinov, Daniel Fried, and Aditi Raghunathan. 2024c · 2024
Cited alongside, same era.
Fangzhou Wu, Ethan Cecchetti, and Chaowei Xiao. 2024b · 2024
Cited alongside, same era.
A new era in llm security: Exploring security concerns in real-world llm-based systems
Fangzhou Wu, Ning Zhang, Somesh Jha, Patrick McDaniel, and Chaowei Xiao. 2024d · 2024
Cited alongside, same era.
Advweb: Controllable black-box attacks on vlm-powered web agents
Chejian Xu, Mintong Kang, Jiawei Zhang, Zeyi Liao, Lingbo Mo, Mengqi Yuan, Huan Sun, and Bo Li. 2024 · 2024
Cited alongside, same era.
Attacking Vision-Language Computer Agents via Pop-ups
Yanzhe Zhang, Tao Yu, and Diyi Yang. 2024 · 2024
Cited alongside, same era.
AWS Identity and Access Management (IAM)
Amazon Web Services. 2025 · 2025
Cited alongside, same era.
Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks
Sizhe Chen, Arman Zharmagambetov, David Wagner, and Chuan Guo. 2025c · 2025
Cited alongside, same era.
How Not to Detect Prompt Injections with an LLM
Sarthak Choudhary, Divyam Anshumaan, Nils Palumbo, and Somesh Jha. 2025 · 2025
Cited alongside, same era.
Llama Prompt Guard 2
Meta. 2025 · 2025
Closest in time.
Azure Policy Documentation
Microsoft. 2025 · 2025
Closest in time.
Adversarial search engine optimization for large language models. In ICLR
Fredrik Nestaas, Edoardo Debenedetti, and Florian Tramèr. 2025 · 2025
Closest in time.
Function calling – OpenAI API
OpenAI. 2025 · 2025
Closest in time.
OpenAI Agents SDK
OpenAI. 2025 · 2025
Closest in time.
python-jsonschema/jsonschema – GitHub
python-jsonschema. 2025 · 2025
Closest in time.
The Dual LLM pattern for building AI assistants that can resist prompt injection
Simon Willison. 2023 · 2025
Closest in time.
Contextual Agent Security: A Policy for Every Purpose. In Proceedings of the 2025 Workshop on Hot Topics in Operating Systems . 8–17
Lillian Tsai and Eugene Bagdasarian. 2025 · 2025
Closest in time.
AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents
Zhun Wang, Vincent Siu, Zhe Ye, Tianneng Shi, Yuzhou Nie, Xuandong Zhao, Chenguang Wang, Wenbo Guo, and Dawn Song. 2025b · 2025
Closest in time.
IsolateGPT: An Execution Isolation Architecture for LLM-Based Systems. In Network and Distributed System Security Symposium (NDSS)
Yuhao Wu, Franziska Roesner, Tadayoshi Kohno, Ning Zhang, and Umar Iqbal. 2025 · 2025
Closest in time.
Agent security bench (ASB): Formalizing and benchmarking attacks and defenses in llm-based agents. In ICLR
Hanrong Zhang, Jingyuan Huang, Kai Mei, Yifei Yao, Zhenting Wang, Chenlu Zhan, Hongwei Wang, and Yongfeng Zhang. 2025 · 2025
Closest in time.