Fetching the paper…
Reading the bibliography…
The emergence of agent-to-agent communication protocols mirrors the early internet: powerful connectivity with minimal security infrastructure.
Privacy as Contextual Integrity
H. Nissenbaum · 2004
Earlier work this paper cites.
The composition theorem for differential privacy
P. Kairouz, S. Oh, and P. Viswanath · 2015
Earlier work this paper cites.
Understanding firewalls for home and small office use
CISA · 2023
Earlier work this paper cites.
Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection
K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz · 2023
Earlier work this paper cites.
Cooperation, competition, and maliciousness: LLM-stakeholders interactive negotiation
S. Abdelnabi, A. Gomaa, S. Sivaprasad, L. Schönherr, and M. Fritz · 2024
Earlier work this paper cites.
AirGapAgent: Protecting Privacy-Conscious Conversational Agents
E. Bagdasarian, R. Yi, S. Ghalebikesabi, P. Kairouz, M. Gruteser, S. Oh, B. Balle, and D. Ramage · 2024
Earlier work this paper cites.
AI Agents with Formal Security Guarantees
M. Balunovic, L. Beurer-Kellner, M. Fischer, and M. Vechev · 2024
Earlier work this paper cites.
Dataset and Lessons Learned from the 2024 SaTML LLM Capture-the-Flag Competition
E. Debenedetti, J. Rando, D. Paleka, S. F. Florin, D. Albastroiu, N. Cohen, Y. Lemberg, R. Ghosh, R. Wen, A. Salem, et al · 2024
Cited alongside, same era.
Operationalizing Contextual Integrity in Privacy-Conscious Assistants
S. Ghalebikesabi, E. Bagdasaryan, R. Yi, I. Yona, I. Shumailov, A. Pappu, C. Shi, L. Weidinger, R. Stanforth, L. Berrada, et al · 2024
Cited alongside, same era.
A scalable communication protocol for networks of large language models
S. Marro, E. La Malfa, J. Wright, G. Li, N. Shadbolt, M. Wooldridge, and P. Torr · 2024
Cited alongside, same era.
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
N. Mireshghallah, H. Kim, X. Zhou, Y. Tsvetkov, M. Sap, R. Shokri, and Y. Choi · 2024
Cited alongside, same era.
Identifying the Risks of LM Agents with an LM-Emulated Sandbox
Y. Ruan, H. Dong, A. Wang, S. Pitis, Y. Zhou, J. Ba, Y. Dubois, C. J. Maddison, and T. Hashimoto · 2024
Cited alongside, same era.
TravelPlanner: A Benchmark for Real-World Planning with Language Agents
J. Xie, K. Zhang, J. Chen, T. Zhu, R. Lou, Y. Tian, Y. Xiao, and Y. Su · 2024
Later among the works it cites.
Get my drift? catching LLM task drift with activation deltas
S. Abdelnabi, A. Fay, G. Cherubin, A. Salem, M. Fritz, and A. Paverd · 2025
Closest in time.
Defeating prompt injections by design
E. Debenedetti, I. Shumailov, T. Fan, J. Hayes, N. Carlini, D. Fabian, C. Kern, C. Shi, A. Terzis, and F. Tramèr · 2025
Closest in time.
Adversarial search engine optimization for large language models
F. Nestaas, E. Debenedetti, and F. Tramèr · 2025
Closest in time.
A.I. Will Empower Humanity
NYT · 2025
Closest in time.
Contextual Agent Security: A Policy for Every Purpose
L. Tsai and E. Bagdasarian · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Privacylens: Evaluating privacy norm awareness of language models in action
Y. Shao, T. Li, W. Shi, Y. Liu, and D. Yang · 2024
Cited alongside, same era.
System-Level Defense against Indirect Prompt Injection Attacks: An Information Flow Control Perspective
F. Wu, E. Cecchetti, and C. Xiao · 2024
Cited alongside, same era.
The Best Chatbot for Hotels with AI
Asksuite
Cited in the paper.
AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents
E. Debenedetti, J. Zhang, M. Balunović, L. Beurer-Kellner, M. Fischer, and F. Tramèr
Cited in the paper.
Chatbots for Reservations and Bookings
FutrAI
Cited in the paper.
MCP Security Notification: Tool Poisoning Attacks
Invariant Labs
Cited in the paper.
Trading Inference-Time Compute for Adversarial Robustness, 2025a
OpenAI
Cited in the paper.
A. Zharmagambetov, C. Guo, I. Evtimov, M. Pavlova, R. Salakhutdinov, and K. Chaudhuri · 2025
Closest in time.