Fetching the paper…
Reading the bibliography…
As AI agents powered by Large Language Models (LLMs) become increasingly versatile and capable of addressing a broad spectrum of tasks, ensuring their security has become a critical challenge.
Mental fatigue: costs and benefits
Maarten AS Boksem and Mattie Tops · 2008
Earlier work this paper cites.
Mapreduce: simplified data processing on large clusters
Jeffrey Dean and Sanjay Ghemawat · 2008
Earlier work this paper cites.
Intriguing properties of neural networks, 2014
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2014
Earlier work this paper cites.
Stop annoying me! an empirical investigation of the usability of app privacy notifications
Nicholas Micallef, Mike Just, Lynne Baillie, and Maher Alharby · 2017
Earlier work this paper cites.
Algorithm appreciation: People prefer algorithmic to human judgment
Jennifer M Logg, Julia A Minson, and Don A Moore · 2019
Earlier work this paper cites.
I genuinely believe prompt engineering is the highest-leverage skill someone can learn in 2022, Sep 2022a
Goodside · 2022
Earlier work this paper cites.
Ignore previous prompt: Attack techniques for language models
Fábio Perez and Ian Ribeiro · 2022
Earlier work this paper cites.
Prompt injection attacks against GPT-3
Simon Willison · 2022
Earlier work this paper cites.
Webshop: Towards scalable real-world web interaction with grounded language agents
Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan · 2022
Earlier work this paper cites.
Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
Sahar Abdelnabi, Kai Greshake, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz · 2023
Earlier work this paper cites.
Hacking Google Bard: From Prompt Injection to Data Exfiltration
Embrace The Red · 2023
Earlier work this paper cites.
Misusing tools in large language models with visual adversarial examples
Xiaohan Fu, Zihan Wang, Shuheng Li, Rajesh K. Gupta, Niloofar Mireshghallah, Taylor Berg-Kirkpatrick, and Earlence Fernandes · 2023
Earlier work this paper cites.
New prompt injection attack on ChatGPT web version. Markdown images can steal your chat data
Samoilenko, Roman · 2023
Earlier work this paper cites.
Delimiters won’t save you from prompt injection
Simon Willison · 2023
Cited alongside, same era.
The Dual LLM Pattern for Building AI Assistants That Can Resist Prompt Injection
Simon Willison · 2023
Cited alongside, same era.
Auto-gpt for online decision making: Benchmarks and additional opinions
Hui Yang, Sifu Yue, and Yunzhong He · 2023
Cited alongside, same era.
Benchmarking and defending against indirect prompt injection attacks on large language models, 2023
Jingwei Yi, Yueqi Xie, Bin Zhu, Emre Kiciman, Guangzhong Sun, Xing Xie, and Fangzhao Wu · 2023
Cited alongside, same era.
Airgapagent: Protecting privacy-conscious conversational agents, 2024
Eugene Bagdasarian, Ren Yi, Sahra Ghalebikesabi, Peter Kairouz, Marco Gruteser, Sewoong Oh, Borja Balle, and Daniel Ramage · 2024
Cited alongside, same era.
Formalizing and Benchmarking Prompt Injection Attacks and Defenses
Yupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia, and Neil Zhenqiang Gong · 2024
Later among the works it cites.
Gorilla: Large language model connected with massive apis
Shishir G. Patil, Tianjun Zhang, Xin Wang, and Joseph E. Gonzalez · 2024
Later among the works it cites.
Fine-tuned deberta-v3-base for prompt injection detection, 2024
ProtectAI.com · 2024
Later among the works it cites.
Exfiltration of personal information from chatgpt via prompt injection, 2024
Gregory Schwartzman · 2024
Later among the works it cites.
Permissive information-flow analysis for large language models, 2024
Shoaib Ahmed Siddiqui, Radhika Gaonkar, Boris Köpf, David Krueger, Andrew Paverd, Ahmed Salem, Shruti Tople, Lukas Wutschitz, Menglin Xia, and Santiago Zanella-Béguelin · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
AI agents with formal security guarantees
Mislav Balunovic, Luca Beurer-Kellner, Marc Fischer, and Martin Vechev · 2024
Cited alongside, same era.
Guiding llms the right way: Fast, non-invasive constrained generation
Luca Beurer-Kellner, Marc Fischer, and Martin T. Vechev · 2024
Cited alongside, same era.
StruQ: Defending against prompt injection with structured queries
Sizhe Chen, Julien Piet, Chawin Sitawarin, and David Wagner · 2024
Cited alongside, same era.
Contextcite: Attributing model generation to context
Benjamin Cohen-Wang, Harshay Shah, Kristian Georgiev, and Aleksander Madry · 2024
Cited alongside, same era.
AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents
Edoardo Debenedetti, Jie Zhang, Mislav Balunović, Luca Beurer-Kellner, Marc Fischer, and Florian Tramèr · 2024
Cited alongside, same era.
ASCII Smuggler Tool: Crafting Invisible Text and Decoding Hidden Codes
Embrace The Red · 2024
Cited alongside, same era.
GitHub Copilot Chat: From Prompt Injection to Data Exfiltration (Copirate)
Embrace The Red · 2024
Cited alongside, same era.
Joseph Spracklen, Raveen Wijewickrama, A H. M. Nazmus Sakib, Anindya Maiti, and Murtuza Jadliwala · 2024
Later among the works it cites.
The instruction hierarchy: Training LLMs to prioritize privileged instructions, 2024
Eric Wallace, Kai Xiao, Reimar Leike, Lilian Weng, Johannes Heidecke, and Alex Beutel · 2024
Later among the works it cites.
Certifiably robust rag against retrieval corruption, 2024
Chong Xiang, Tong Wu, Zexuan Zhong, David Wagner, Danqi Chen, and Prateek Mittal · 2024
Later among the works it cites.
Swe-agent: Agent-computer interfaces enable automated software engineering
John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press · 2024
Later among the works it cites.
Improving alignment and robustness with circuit breakers
Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin, Maksym Andriushchenko, J. Zico Kolter, Matt Fredrikson, and Dan Hendrycks · 2024
Later among the works it cites.
Can LLMs separate instructions from data? and what do we even mean by that?
Egor Zverev, Sahar Abdelnabi, Mario Fritz, and Christoph H Lampert · 2024
Later among the works it cites.
Defeating prompt injections by design, 2025
Edoardo Debenedetti, Ilia Shumailov, Tianqi Fan, Jamie Hayes, Nicholas Carlini, Daniel Fabian, Christoph Kern, Chongyang Shi, Andreas Terzis, and Florian Tramèr · 2025
Closest in time.
IsolateGPT: An Execution Isolation Architecture for LLM-Based Agentic Systems
Yuhao Wu, Franziska Roesner, Tadayoshi Kohno, Ning Zhang, and Umar Iqbal · 2025
Closest in time.