Fetching the paper…
Reading the bibliography…
As AI agents become increasingly autonomous and capable, ensuring their security against vulnerabilities such as prompt injection becomes critical.
A lattice model of secure information flow
Dorothy E Denning · 1976
Earlier work this paper cites.
Security policies and security models
Joseph A Goguen and José Meseguer · 1982
Earlier work this paper cites.
A decentralized model for information flow control
Andrew C. Myers and Barbara Liskov · 1997
Earlier work this paper cites.
Safety versus secrecy
Dennis Volpano · 1999
Earlier work this paper cites.
Language-based information-flow security
Andrei Sabelfeld and Andrew C Myers · 2003
Earlier work this paper cites.
Enforcing robust declassification
Andrew C Myers, Andrei Sabelfeld, and Steve Zdancewic · 2004
Earlier work this paper cites.
Explicit secrecy: A policy for taint tracking
Daniel Schoepe, Musard Balliu, Benjamin C. Pierce, and Andrei Sabelfeld · 2016
Earlier work this paper cites.
LangChain, 2022
Harrison Chase · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe · 2022
Earlier work this paper cites.
Grammar-constrained decoding for structured nlp tasks without finetuning
Saibo Geng, Martin Josifoski, Maxime Peyrard, and Robert West · 2023
Earlier work this paper cites.
Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection, 2023
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz · 2023
Earlier work this paper cites.
Prompt injection attack against LLM-integrated applications, 2023
Yi Liu, Gelei Deng, Yuekang Li, Kailong Wang, Zihao Wang, Xiaofeng Wang, Tianwei Zhang, Yepang Liu, Haoyu Wang, Yan Zheng, and Yang Liu · 2023
Earlier work this paper cites.
Toolformer: Language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessi, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom · 2023
Earlier work this paper cites.
The dual LLM pattern for building ai assistants that can resist prompt injection
Simon Willison · 2023
Earlier work this paper cites.
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao · 2023
Cited alongside, same era.
Benchmarking and defending against indirect prompt injection attacks on large language models, 2023
Jingwei Yi, Yueqi Xie, Bin Zhu, Emre Kiciman, Guangzhong Sun, Xing Xie, and Fangzhao Wu · 2023
Cited alongside, same era.
Computer Use (beta)
Anthropic · 2024
Cited alongside, same era.
Embedding-based classifiers can detect prompt injection attacks
Md. Ahsan Ayub and Subhabrata Majumdar · 2024
Cited alongside, same era.
AI agents with formal security guarantees
Mislav Balunovic, Luca Beurer-Kellner, Marc Fischer, and Martin Vechev · 2024
Cited alongside, same era.
Guiding llms the right way: Fast, non-invasive constrained generation, 2024
Luca Beurer-Kellner, Marc Fischer, and Martin Vechev · 2024
The instruction hierarchy: Training LLMs to prioritize privileged instructions, 2024
Eric Wallace, Kai Xiao, Reimar Leike, Lilian Weng, Johannes Heidecke, and Alex Beutel · 2024
Later among the works it cites.
System-level defense against indirect prompt injection attacks: An information flow control perspective, 2024
Fangzhou Wu, Ethan Cecchetti, and Chaowei Xiao · 2024
Later among the works it cites.
AutoGen: Enabling next-gen LLM applications via multi-agent conversations
Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, Ahmed Hassan Awadallah, Ryen W White, Doug Burger, and Chi Wang · 2024
Later among the works it cites.
GradSafe: Detecting jailbreak prompts for LLMs via safety-critical gradient analysis
Yueqi Xie, Minghong Fang, Renjie Pi, and Neil Gong · 2024
Later among the works it cites.
Injecagent: Benchmarking indirect prompt injections in tool-integrated large language model agents
Qiusi Zhan, Zhixiang Liang, Zifan Ying, and Daniel Kang · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Agentdojo: A dynamic environment to evaluate prompt injection attacks and defenses for llm agents, 2024
Edoardo Debenedetti, Jie Zhang, Mislav Balunović, Luca Beurer-Kellner, Marc Fischer, and Florian Tramèr · 2024
Cited alongside, same era.
Magentic-One: A generalist multi-agent system for solving complex tasks, 2024
Adam Fourney, Gagan Bansal, Hussein Mozannar, Cheng Tan, Eduardo Salinas, Erkang, Zhu, Friederike Niedtner, Grace Proebsting, Griffin Bassman, Jack Gerrits, Jacob Alber, Peter Chang, Ricky Loynd, Robert West, Victor Dibia, Ahmed Awadallah, Ece Kamar, Rafah Hosn, and Saleema Amershi · 2024
Cited alongside, same era.
Defending against indirect prompt injection attacks with spotlighting
Keegan Hines, Gary Lopez, Matthew Hall, Federico Zarfati, Yonatan Zunger, and Emre Kiciman · 2024
Cited alongside, same era.
The task shield: Enforcing task alignment to defend against indirect prompt injection in llm agents, 2024
Feiran Jia, Tong Wu, Xin Qin, and Anna Squicciarini · 2024
Cited alongside, same era.
Devin: The first AI software engineer
Cognition Labs · 2024
Cited alongside, same era.
Formalizing and benchmarking prompt injection attacks and defenses
Yupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia, and Neil Zhenqiang Gong · 2024
Cited alongside, same era.
Improving alignment and robustness with circuit breakers
Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin, Maksym Andriushchenko, Rowan Wang, Zico Kolter, Matt Fredrikson, and Dan Hendrycks · 2024
Later among the works it cites.
Get my drift? catching llm task drift with activation deltas
Sahar Abdelnabi, Aideen Fay, Giovanni Cherubin, Ahmed Salem, and Mario Fritz · 2025
Closest in time.
Guidance: A guidance language for controlling large language models
Guidance AI · 2025
Closest in time.
StruQ: Defending against prompt injection with structured queries
Sizhe Chen, Julien Piet, Chawin Sitawarin, and David Wagner · 2025
Closest in time.
Secalign: Defending against prompt injection with preference optimization, 2025
Sizhe Chen, Arman Zharmagambetov, Saeed Mahloujifar, Kamalika Chaudhuri, David Wagner, and Chuan Guo · 2025
Closest in time.
Defeating prompt injections by design, 2025
Edoardo Debenedetti, Ilia Shumailov, Tianqi Fan, Jamie Hayes, Nicholas Carlini, Daniel Fabian, Christoph Kern, Chongyang Shi, Andreas Terzis, and Florian Tramèr · 2025
Closest in time.
Instructional segment embedding: Improving LLM safety with instruction hierarchy
Tong Wu, Shujian Zhang, Kaiqiang Song, Silei Xu, Sanqiang Zhao, Ravi Agrawal, Sathish Reddy Indurthi, Chong Xiang, Prateek Mittal, and Wenxuan Zhou · 2025
Closest in time.
Agent security bench (asb): Formalizing and benchmarking attacks and defenses in llm-based agents
Hanrong Zhang, Jingyuan Huang, Kai Mei, Yifei Yao, Zhenting Wang, Chenlu Zhan, Hongwei Wang, and Yongfeng Zhang · 2025
Closest in time.
Rtbas: Defending llm agents against prompt injection and privacy leakage, 2025
Peter Yong Zhong, Siyuan Chen, Ruiqi Wang, McKenna McCall, Ben L. Titzer, Heather Miller, and Phillip B. Gibbons · 2025
Closest in time.