Fetching the paper…
Reading the bibliography…
Gemini is increasingly used to perform tasks on behalf of users, where function-calling and tool-use capabilities enable the model to access user data.
Evasion attacks against machine learning at test time
B. Biggio, I. Corona, D. Maiorca, B. Nelson, N. Šrndić, P. Laskov, G. Giacinto, and F. Roli · 2013
Earlier work this paper cites.
Fundamental limits on adversarial robustness
A. Fawzi, O. Fawzi, and P. Frossard · 2015
Earlier work this paper cites.
Explaining and harnessing adversarial examples, 2015
I. J. Goodfellow, J. Shlens, and C. Szegedy · 2015
Earlier work this paper cites.
Transferability in machine learning: From phenomena to black-box attacks using adversarial samples
N. Papernot, P. McDaniel, and I. Goodfellow · 2016
Earlier work this paper cites.
Universal adversarial perturbations
S.-M. Moosavi-Dezfooli, A. Fawzi, O. Fawzi, and P. Frossard · 2017
Earlier work this paper cites.
J. Gilmer, L. Metz, F. Faghri, S. S. Schoenholz, M. Raghu, M. Wattenberg, and I. Goodfellow · 2018
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu · 2018
Earlier work this paper cites.
On evaluating adversarial robustness
N. Carlini, A. Athalye, N. Papernot, W. Brendel, J. Rauber, D. Tsipras, I. Goodfellow, A. Madry, and A. Kurakin · 2019
Earlier work this paper cites.
Adversarial examples are not bugs, they are features
A. Ilyas, S. Santurkar, D. Tsipras, L. Engstrom, B. Tran, and A. Madry · 2019
Earlier work this paper cites.
Robustness may be at odds with accuracy
D. Tsipras, S. Santurkar, L. Engstrom, A. Turner, and A. Madry · 2019
Earlier work this paper cites.
On adaptive attacks to adversarial example defenses
F. Tramèr, N. Carlini, W. Brendel, and A. Mądry · 2020
Earlier work this paper cites.
On the limitations of stochastic pre-processing defenses
Y. Gao, I. Shumailov, K. Fawaz, and N. Papernot · 2022
Earlier work this paper cites.
Robustness and accuracy could be reconcilable by (proper) definition
T. Pang, M. Lin, X. Yang, J. Zhu, and S. Yan · 2022
Earlier work this paper cites.
Google ai principles, 2023
Google · 2023
Earlier work this paper cites.
Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection
K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz · 2023
Earlier work this paper cites.
Baseline defenses for adversarial attacks against aligned language models, 2023
N. Jain, A. Schwarzschild, Y. Wen, G. Somepalli, J. Kirchenbauer, P. yeh Chiang, M. Goldblum, A. Saha, J. Geiping, and T. Goldstein · 2023
Earlier work this paper cites.
Sok: Certified robustness for deep neural networks
L. Li, T. Xie, and B. Li · 2023
Earlier work this paper cites.
Prompt injection attack against llm-integrated applications, 2023
Y. Liu, G. Deng, Y. Li, K. Wang, T. Zhang, Y. Liu, H. Wang, Y. Zheng, and Y. Liu · 2023
Earlier work this paper cites.
Randomness in ml defenses helps persistent attackers and hinders evaluators, 2023
K. Lucas, M. Jagielski, F. Tramèr, L. Bauer, and N. Carlini · 2023
Earlier work this paper cites.
New prompt injection attack on chatgpt web version
R. Samoilenko · 2023
Earlier work this paper cites.
Benchmarking and defending against indirect prompt injection attacks on large language models
J. Yi, Y. Xie, B. Zhu, K. Hines, E. Kiciman, G. Sun, X. Xie, and F. Wu · 2023
Earlier work this paper cites.
Universal and transferable adversarial attacks on aligned language models
A. Zou, Z. Wang, N. Carlini, M. Nasr, J. Z. Kolter, and M. Fredrikson · 2023
Earlier work this paper cites.
Jailbreaking leading safety-aligned LLMs with simple adaptive attacks, Apr. 2024
M. Andriushchenko, F. Croce, and N. Flammarion · 2024
Cited alongside, same era.
Defending against unforeseen failure modes with latent adversarial training, 2024
S. Casper, L. Schulze, O. Patel, and D. Hadfield-Menell · 2024
Cited alongside, same era.
StruQ: Defending against prompt injection with structured queries, Feb. 2024
S. Chen, J. Piet, C. Sitawarin, and D. Wagner · 2024
Cited alongside, same era.
Agentdojo: A dynamic environment to evaluate prompt injection attacks and defenses for LLM agents
E. Debenedetti, J. Zhang, M. Balunovic, L. Beurer-Kellner, M. Fischer, and F. Tramèr · 2024
Cited alongside, same era.
Imprompter: Tricking llm agents into improper tool use
X. Fu, S. Li, Z. Wang, Y. Liu, R. K. Gupta, T. Berg-Kirkpatrick, and E. Fernandes · 2024
The instruction hierarchy: Training llms to prioritize privileged instructions, 2024
E. Wallace, K. Xiao, R. Leike, L. Weng, J. Heidecke, and A. Beutel · 2024
Later among the works it cites.
Jailbreak and guard aligned language models with only few in-context demonstrations, 2024
Z. Wei, Y. Wang, A. Li, Y. Mo, and Y. Wang · 2024
Later among the works it cites.
Efficient adversarial training in llms with continuous attacks, 2024
S. Xhonneux, A. Sordoni, S. Günnemann, G. Gidel, and L. Schwinn · 2024
Later among the works it cites.
Robust llm safeguarding via refusal feature adversarial training
L. Yu, V. Do, K. Hambardzumyan, and N. Cancedda · 2024
Later among the works it cites.
Air-bench 2024: A safety benchmark based on risk categories from regulations and policies, 2024b
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Defending against indirect prompt injection attacks with spotlighting, 2024
K. Hines, G. Lopez, M. Hall, F. Zarfati, Y. Zunger, and E. Kiciman · 2024
Cited alongside, same era.
J. Hughes, S. Price, A. Lynch, R. Schaeffer, F. Barez, S. Koyejo, H. Sleight, E. Jones, E. Perez, and M. Sharma · 2024
Cited alongside, same era.
Attention tracker: Detecting prompt injection attacks in llms, 2024
K.-H. Hung, C.-Y. Ko, A. Rawat, I.-H. Chung, W. H. Hsu, and P.-Y. Chen · 2024
Cited alongside, same era.
Robust safety classifier against jailbreaking attacks: Adversarial prompt shield
J. Kim, A. Derakhshan, and I. Harris · 2024
Cited alongside, same era.
St-webagentbench: A benchmark for evaluating safety and trustworthiness in web agents, 2024
I. Levy, B. Wiesel, S. Marreed, A. Oved, A. Yaeli, and S. Shlomov · 2024
Cited alongside, same era.
Adversarial tuning: Defending against jailbreak attacks for llms
F. Liu, Z. Xu, and H. Liu · 2024
Cited alongside, same era.
New gemini for workspace vulnerability enabling phishing & content manipulation
J. Martin and K. Yeung · 2024
Cited alongside, same era.
Y. Zeng, Y. Yang, A. Zhou, J. Z. Tan, Y. Tu, Y. Mai, K. Klyman, M. Pan, R. Jia, D. Song, P. Liang, and B. Li · 2024
Later among the works it cites.
InjecAgent: Benchmarking indirect prompt injections in tool-integrated large language model agents
Q. Zhan, Z. Liang, Z. Ying, and D. Kang · 2024
Later among the works it cites.
Improving alignment and robustness with circuit breakers, June 2024
A. Zou, L. Phan, J. Wang, D. Duenas, M. Lin, M. Andriushchenko, R. Wang, Z. Kolter, M. Fredrikson, and D. Hendrycks · 2024
Later among the works it cites.
Get my drift? catching llm task drift with activation deltas, 2025
S. Abdelnabi, A. Fay, G. Cherubin, A. Salem, M. Fritz, and A. Paverd · 2025
Closest in time.
Agentharm: A benchmark for measuring harmfulness of llm agents, 2025
M. Andriushchenko, A. Souly, M. Dziemian, D. Duenas, M. Lin, J. Wang, D. Hendrycks, A. Zou, Z. Kolter, M. Fredrikson, E. Winsor, J. Wynne, Y. Gal, and X. Davies · 2025
Closest in time.
Doomarena: A framework for testing ai agents against evolving security threats
L. Boisvert, M. Bansal, C. K. R. Evuru, G. Huang, A. Puri, A. Bose, M. Fazel, Q. Cappart, J. Stanley, A. Lacoste, et al · 2025
Closest in time.
SecAlign: Defending against prompt injection with preference optimization, Jan. 2025
S. Chen, A. Zharmagambetov, S. Mahloujifar, K. Chaudhuri, D. Wagner, and C. Guo · 2025
Closest in time.
Llamafirewall: An open source guardrail system for building secure ai agents, 2025
S. Chennabasappa, C. Nikolaidis, D. Song, D. Molnar, S. Ding, S. Wan, S. Whitman, L. Deason, N. Doucette, A. Montilla, A. Gampa, B. de Paola, D. Gabi, J. Crnkovich, J.-C. Testud, K. He, R. Chaturvedi, W. Zhou, and J. Saxe · 2025
Closest in time.
Defeating prompt injections by design, 2025
E. Debenedetti, I. Shumailov, T. Fan, J. Hayes, N. Carlini, D. Fabian, C. Kern, C. Shi, A. Terzis, and F. Tramèr · 2025
Closest in time.
Function calling with the gemini api, 2025
Google · 2025
Closest in time.
A. Labunets, N. V. Pandya, A. Hooda, X. Fu, and E. Fernandes · 2025
Closest in time.
DataSentinel: A game-theoretic detection of prompt injection attacks
Y. Liu, Y. Jia, J. Jia, D. Song, and N. Z. Gong · 2025
Closest in time.
Adversarial training for multimodal large language models against jailbreak attacks
L. Lu, S. Pang, S. Liang, H. Zhu, X. Zeng, A. Liu, Y. Liu, and Y. Zhou · 2025
Closest in time.
Hacking gemini’s memory with prompt injection and delayed tool invocation
J. Rehberger · 2025
Closest in time.
Safearena: Evaluating the safety of autonomous web agents, 2025
A. D. Tur, N. Meade, X. H. Lù, A. Zambrano, A. Patel, E. Durmus, S. Gella, K. Stańczak, and S. Reddy · 2025
Closest in time.
Instructional segment embedding: Improving LLM safety with instruction hierarchy
T. Wu, S. Zhang, K. Song, S. Xu, S. Zhao, R. Agrawal, S. R. Indurthi, C. Xiang, P. Mittal, and W. Zhou · 2025
Closest in time.
Trading inference-time compute for adversarial robustness, 2025
W. Zaremba, E. Nitishinskaya, B. Barak, S. Lin, S. Toyer, Y. Yu, R. Dias, E. Wallace, K. Xiao, J. Heidecke, and A. Glaese · 2025
Closest in time.