Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are increasingly deployed in agentic systems that interact with an untrusted environment.
“A lattice model of secure information flow”
Dorothy Denning · 1976
Earlier work this paper cites.
“Certification of programs for secure information flow”
Dorothy. Denning and Peter. Denning · 1977
Earlier work this paper cites.
“The Cambridge CAP computer and its protection system”
Roger Needham and Robin Walker · 1977
Earlier work this paper cites.
“Intelligent agents: Theory and practice”
Michael Wooldridge and Nicholas Jennings · 1995
Earlier work this paper cites.
“Smashing The Stack For Fun And Profit”
Aleph One · 1996
Earlier work this paper cites.
“A decentralized model for information flow control”
Andrew. Myers and Barbara Liskov · 1997
Earlier work this paper cites.
“Security policies”
Ross Anderson, Frank Stajano and Jong-Hyeon Lee · 2002
Earlier work this paper cites.
“Language-based information-flow security”
Andrei Sabelfeld and Andrew Myers · 2003
Earlier work this paper cites.
“The geometry of innocent flesh on the bone: return-into-libc without function calls (on the x86)”
Hovav Shacham · 2007
Earlier work this paper cites.
“Control-flow integrity principles, implementations, and applications”
Martín Abadi, Mihai Budiu, Ulfar Erlingsson and Jay Ligatti · 2009
Earlier work this paper cites.
“Security engineering: a guide to building dependable distributed systems”, 2010
Ross Anderson · 2010
Earlier work this paper cites.
“Capsicum: Practical Capabilities for UNIX”
Robert Watson, Jonathan Anderson, Ben Laurie and Kris Kennaway · 2010
Earlier work this paper cites.
“Android permissions: user attention, comprehension, and behavior”
Adrienne Felt, Elizabeth Ha, Serge Egelman, Ariel Haney, Erika Chin and David Wagner · 2012
Earlier work this paper cites.
“libcap: POSIX capabilities support for Linux”, 2013
Andrew. Morgan · 2013
Earlier work this paper cites.
“ROP is still dangerous: Breaking modern defenses”
Nicholas Carlini and David Wagner · 2014
Earlier work this paper cites.
“The CHERI capability model: Revisiting RISC in an age of risk”
Jonathan Woodruff, Robert Watson, David Chisnall, Simon Moore, Jonathan Anderson, Brooks Davis, Ben Laurie, Peter Neumann, Robert Norton and Michael Roe · 2014
Earlier work this paper cites.
“Clean Application Compartmentalization with SOAAP”
Khilan Gudka, Robert.M. Watson, Jonathan Anderson, David Chisnall, Brooks Davis, Ben Laurie, Ilias Marinos, Peter. Neumann and Alex Richardson · 2015
Earlier work this paper cites.
“CHERI: A hybrid capability-system architecture for scalable software compartmentalization”
Robert Watson, Jonathan Woodruff, Peter Neumann, Simon Moore, Jonathan Anderson, David Chisnall, Nirav Dave, Brooks Davis, Khilan Gudka and Ben Laurie · 2015
Earlier work this paper cites.
“Audit Committee update: Insider Threat”, 2018
PricewaterhouseCoopers · 2018
Earlier work this paper cites.
“A Large Scale Study of User Behavior, Expectations and Engagement with Android Permissions”
Weicheng Cao, Chunqiu Xia, Sai Peddinti, David Lie, Nina Taft and Lisa. Austin · 2021
Earlier work this paper cites.
“WebGPT: Browser-assisted question-answering with human feedback”
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju and William Saunders · 2021
Earlier work this paper cites.
“Exploiting GPT-3 prompts with malicious inputs that order the model to ignore its previous directions”, 2022
Riley Goodside · 2022
Earlier work this paper cites.
“tiktoken: Fast BPE tokeniser for use with OpenAI’s models”
OpenAI · 2022
Earlier work this paper cites.
“Ignore previous prompt: Attack techniques for language models”
Fábio Perez and Ian Ribeiro · 2022
Earlier work this paper cites.
“Cost Of Insider Threats Global Report”, 2022
Ponemon-Institute · 2022
Earlier work this paper cites.
“LaMDA: Language models for dialog applications”
Romal Thoppilan, Daniel De, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker and Yu Du · 2022
Cited alongside, same era.
“ReAct: Synergizing reasoning and acting in language models”
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan and Yuan Cao · 2022
Cited alongside, same era.
“Are aligned neural networks adversarially aligned?”
Nicholas Carlini, Milad Nasr, Christopher. Choquette-Choo, Matthew Jagielski, Irena Gao, Pang Koh, Daphne Ippolito, Florian Tramèr and Ludwig Schmidt · 2023
Cited alongside, same era.
“PAL: Program-aided language models”
Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan and Graham Neubig · 2023
Cited alongside, same era.
“Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection”
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz and Mario Fritz · 2023
“Fine-Tuned DeBERTa-v3-base for Prompt Injection Detection”
ProtectAI · 2024
Later among the works it cites.
“Embrace The Red Blog”, 2024
Johann Rehberger · 2024
Later among the works it cites.
“SPML: A DSL for Defending Language Models Against Prompt Attacks”
Reshabh Sharma, Vinayak Gupta and Dan Grossman · 2024
Later among the works it cites.
“HuggingGPT: Solving AI tasks with ChatGPT and its friends in Hugging Face”
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu and Yueting Zhuang · 2024
Later among the works it cites.
“The instruction hierarchy: Training llms to prioritize privileged instructions”
Eric Wallace, Kai Xiao, Reimar Leike, Lilian Weng, Johannes Heidecke and Alex Beutel · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Communication-Efficient Learning of Deep Networks from Decentralized Data”, 2023
H. McMahan, Eider Moore, Daniel Ramage, Seth Hampson and Blaiseüera y Arcas · 2023
Cited alongside, same era.
“ToolLLM: Facilitating large language models to master 16000+ real-world APIs”
Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang and Bill Qian · 2023
Cited alongside, same era.
“ToolFormer: Language Models Can Teach Themselves to Use Tools”
Timo Schick, Jane Dwivedi-Yu, Roberto Dessi, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda and Thomas Scialom · 2023
Cited alongside, same era.
“The Dual LLM pattern for building AI assistants that can resist prompt injection”, 2023
Simon Willison · 2023
Cited alongside, same era.
“Are you still on track!? Catching LLM Task Drift with Activations”
Sahar Abdelnabi, Aideen Fay, Giovanni Cherubin, Ahmed Salem, Mario Fritz and Andrew Paverd · 2024
Cited alongside, same era.
“The Claude 3 Model Family: Opus, Sonnet, Haiku”, 2024
Anthropic · 2024
Cited alongside, same era.
“Air Gap: Protecting Privacy-Conscious Conversational Agents”
Eugene Bagdasaryan, Ren Yi, Sahra Ghalebikesabi, Peter Kairouz, Marco Gruteser, Sewoong Oh, Borja Balle and Daniel Ramage · 2024
Cited alongside, same era.
Weizhi Wang, Li Dong, Hao Cheng, Xiaodong Liu, Xifeng Yan, Jianfeng Gao and Furu Wei · 2024
Later among the works it cites.
Fangzhou Wu, Ethan Cecchetti and Chaowei Xiao · 2024
Later among the works it cites.
“Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy”
Tong Wu, Shujian Zhang, Kaiqiang Song, Silei Xu, Sanqiang Zhao, Ravi Agrawal, Sathish Indurthi, Chong Xiang, Prateek Mittal and Wenxuan Zhou · 2024
Later among the works it cites.
“Formal Mechanised Semantics of CHERI C: Capabilities, Undefined Behaviour, and Provenance”
Vadim Zaliva, Kayvan Memarian, Ricardo Almeida, Jessica Clarke, Brooks Davis, Alexander Richardson, David Chisnall, Brian Campbell, Ian Stark, Robert.. Watson and Peter Sewell · 2024
Later among the works it cites.
“Firewalls to Secure Dynamic LLM Agentic Networks”
Sahar Abdelnabi, Amr Gomaa, Eugene Bagdasarian, Per Kristensson and Reza Shokri · 2025
Closest in time.
“Technical Blog: Strengthening AI Agent Hijacking Evaluations”, 2025
US-AISI · 2025
Closest in time.
“Mitigate jailbreaks and prompt injections”, 2025
Anthropic · 2025
Closest in time.
“Pydantic”, 2025
Samuel Colvin, Eric Jolibois, Hasan Ramezani, Adrian Garcia, Terrence Dorsey, David Montague, Serge Matveenko, Marcelo Trylesinski, Sydney Runkle, David Hewitt, Alex Hall and Victorien Plot · 2025
Closest in time.
“Securing AI Agents with Information-Flow Control”
Manuel Costa, Boris Köpf, Aashish Kolluri, Andrew Paverd, Mark Russinovich, Ahmed Salem, Shruti Tople, Lukas Wutschitz and Santiago Zanella-Béguelin · 2025
Closest in time.
“Breach By A Thousand Leaks: Unsafe Information Leakage in “Safe” AI Responses”
David Glukhov, Ziwen Han, Ilia Shumailov, Vardan Papyan and Nicolas Papernot · 2025
Closest in time.
“Prompt flow integrity to prevent privilege escalation in llm agents”
Juhee Kim, Woohyuk Choi and Byoungyoung Lee · 2025
Closest in time.
“ACE: A Security Architecture for LLM-Integrated App Systems”
Evan Li, Tushin Mallick, Evan Rose, William Robertson, Alina Oprea and Cristina Nita-Rotaru · 2025
Closest in time.
“Adversarial Search Engine Optimization for Large Language Models”
Fredrik Nestaas, Edoardo Debenedetti and Florian Tramèr · 2025
Closest in time.
“Ignore untrusted data by default”, 2025
OpenAI · 2025
Closest in time.
“CHERI adoption and diffusion research”, 2025
RSM UK Consulting LLP for UK DSIT · 2025
Closest in time.
“How we estimate the risk from prompt injection attacks on AI systems”, 2025
AI-Security-Team, Aneesh Pappu, Andreas Terzis, Chongyang Shi, Gena Gibson, Ilia Shumailov, Itay Yona, Jamie Hayes, John Flynn, Juliette Pluto, Sharon Lin and Shuang Song · 2025
Closest in time.
“Progent: Programmable Privilege Control for LLM Agents”
Tianneng Shi, Jingxuan He, Zhun Wang, Linyu Wu, Hongwei Li, Wenbo Guo and Dawn Song · 2025
Closest in time.
“IsolateGPT: An Execution Isolation Architecture for LLM-Based Agentic Systems”
Yuhao Wu, Franziska Roesner, Tadayoshi Kohno, Ning Zhang and Umar Iqbal · 2025
Closest in time.
“RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage”
Peter Zhong, Siyuan Chen, Ruiqi Wang, McKenna McCall, Ben Titzer and Heather Miller · 2025
Closest in time.
“ASIDE: Architectural Separation of Instructions and Data in Language Models”
Egor Zverev, Evgenii Kortukov, Alexander Panfilov, Alexandra Volkova, Soroush Tabesh, Sebastian Lapuschkin, Wojciech Samek and Christoph Lampert · 2025
Closest in time.