Fetching the paper…
Reading the bibliography…
Current evaluations of defenses against prompt attacks in large language model (LLM) applications often overlook two critical factors: the dynamic nature of adversarial behavior and the usability penalties imposed on legitimate users by restrictive defenses.
Active learning with statistical models
Cohn, D. A., Ghahramani, Z., and Jordan, M. I · 1996
Earlier work this paper cites.
A classification of sql injection attacks and countermeasures
Halfond, W. G., Viegas, J., and Orso, A · 2006
Earlier work this paper cites.
Causality
Pearl, J · 2009
Earlier work this paper cites.
Usable security: History, themes, and challenges
Garfinkel, S. and Lipford, H. R · 2014
Earlier work this paper cites.
The art of cybersecurity: Defense in depth strategy for robust protection
Mughal, A. A · 2018
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
Constitutional AI: Harmlessness from AI feedback
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., Chen, C., Olsson, C., Olah, C., Hernandez, D., Drain, D., Ganguli, D., Li, D., Tran-Johnson, E., Perez, E., Kerr, J., Mueller, J., Ladish, J., Landau, J., Ndousse, K., Lukosuite, K., Lovitt, L., Sellitto, M., Elhage, N., Schiefer, N., Mercado, N., DasSarma, N., Lasenby, R., Larson, R., Ringer, S., Johnston, S., Kravec, S., El Showk, S., Fort, S., Lanham, T., Telleen-Lawton, T., Conerly, T., Henighan, T., Hume, T., Bowman, S. R., Hatfield-Dodds, Z., Mann, B., Amodei, D., Joseph, N., McCandlish, S., Brown, T., and Kaplan, J · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., Jones, A., Bowman, S., Chen, A., Conerly, T., DasSarma, N., Drain, D., Elhage, N., El-Showk, S., Fort, S., Hatfield-Dodds, Z., Henighan, T., Hernandez, D., Hume, T., Jacobson, J., Johnston, S., Kravec, S., Olsson, C., Ringer, S., Tran-Johnson, E., Amodei, D., Brown, T., Joseph, N., McCandlish, S., Olah, C., Kaplan, J., and Clark, J · 2022
Earlier work this paper cites.
harmbenchuage models with language models
Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., and Irving, G · 2022
Earlier work this paper cites.
Attack prompt generation for red teaming and defending large language models
Deng, B., Wang, W., Feng, F., Deng, Y., Wang, Q., and He, X · 2023
Earlier work this paper cites.
Enhancing chat language models by scaling high-quality instructional conversations
Ding, N., Chen, Y., Xu, B., Qin, Y., Hu, S., Liu, Z., Sun, M., and Zhou, B · 2023
Earlier work this paper cites.
Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection
Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., and Fritz, M · 2023
Earlier work this paper cites.
Exploiting programmatic behavior of LLMs: Dual-use through standard security attacks
Kang, D., Li, X., Stoica, I., Guestrin, C., Zaharia, M., and Hashimoto, T · 2023
Earlier work this paper cites.
Ignore this title and HackAPrompt: Exposing systemic vulnerabilities of LLMs through a global prompt hacking competition
Schulhoff, S., Pinto, J., Khan, A., Bouchard, L.-F., Si, C., Anati, S., Tagliabue, V., Kost, A., Carnahan, C., and Boyd-Graber, J · 2023
Earlier work this paper cites.
Multi-agent collaboration: Harnessing the power of intelligent LLM agents
Talebirad, Y. and Nadiri, A · 2023
Cited alongside, same era.
Tensor trust: Interpretable prompt injection attacks from an online game
Toyer, S., Watkins, O., Mendes, E. A., Svegliato, J., Bailey, L., Wang, T., Ong, I., Elmaaroufi, K., Abbeel, P., Darrell, T., Ritter, A., and Russell, S · 2023
Cited alongside, same era.
Jailbreak and guard aligned language models with only few in-context demonstrations
Wei, Z., Wang, Y., Li, A., Mo, Y., and Wang, Y · 2023
Cited alongside, same era.
Benchmarking and defending against indirect prompt injection attacks on large language models
Yi, J., Xie, Y., Zhu, B., Kiciman, E., Sun, G., Xie, X., and Wu, F · 2023
Cited alongside, same era.
Low-resource languages jailbreak GPT-4
Yong, Z. X., Menghini, C., and Bach, S · 2023
Sandwich defense, 2024
Learn Prompting · 2024
Later among the works it cites.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., Sakhaee, E., Li, N., Basart, S., Li, B., Forsyth, D., and Hendrycks, D · 2024
Later among the works it cites.
Neural Exec: Learning (and learning from) execution triggers for prompt injection attacks
Pasquini, D., Strohmeier, M., and Troncoso, C · 2024
Later among the works it cites.
Jatmo: Prompt injection defense by task-specific finetuning
Piet, J., Alrashed, M., Sitawarin, C., Chen, S., Wei, Z., Sun, E., Alomair, B., and Wagner, D · 2024
Later among the works it cites.
An early categorization of prompt injection attacks on large language models
Rossi, S., Michel, A. M., Mukkamala, R. R., and Thatcher, J. B · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., and Fredrikson, M · 2023
Cited alongside, same era.
Embedding-based classifiers can detect prompt injection attacks
Ayub, M. A. and Majumdar, S · 2024
Cited alongside, same era.
Jailbreakbench: An open robustness benchmark for jailbreaking large language models
Chao, P., Debenedetti, E., Robey, A., Andriushchenko, M., Croce, F., Sehwag, V., Dobriban, E., Flammarion, N., Pappas, G. J., Tramèr, F., Hassani, H., and Wong, E · 2024
Cited alongside, same era.
AgentDojo: A dynamic environment to evaluate prompt injection attacks and defenses for LLM agents
Debenedetti, E., Zhang, J., Balunovic, M., Beurer-Kellner, L., Fischer, M., and Tramèr, F · 2024
Cited alongside, same era.
Defending against indirect prompt injection attacks with spotlighting
Hines, K., Lopez, G., Hall, M., Zarfati, F., Zunger, Y., and Kiciman, E · 2024
Cited alongside, same era.
ArtPrompt: ASCII art-based jailbreak attacks against aligned LLMs
Jiang, F., Xu, Z., Niu, L., Xiang, Z., Ramasubramanian, B., Li, B., and Poovendran, R · 2024
Cited alongside, same era.
Robust safety classifier against jailbreaking attacks: Adversarial prompt shield
Kim, J., Derakhshan, A., and Harris, I · 2024
Cited alongside, same era.
Sawtell, M., Masterman, T., Besen, S., and Brown, J · 2024
Later among the works it cites.
Operationalizing a threat model for red-teaming large language models (LLMs)
Verma, A., Krishna, S., Gehrmann, S., Seshadri, M., Pradhan, A., Ault, T., Barrett, L., Rabinowitz, D., Doucette, J., and Phan, N · 2024
Later among the works it cites.
The instruction hierarchy: Training LLMs to prioritize privileged instructions
Wallace, E., Xiao, K., Leike, R., Weng, L., Heidecke, J., and Beutel, A · 2024
Later among the works it cites.
How johnny can persuade LLMs to jailbreak them: Rethinking persuasion to challenge AI safety by humanizing LLMs
Zeng, Y., Lin, H., Zhang, J., Yang, D., Jia, R., and Shi, W · 2024
Later among the works it cites.
Robust prompt optimization for defending language models against jailbreaking attacks
Zhou, A., Li, B., and Wang, H · 2024
Later among the works it cites.
Promptbench: A unified library for evaluation of large language models
Zhu, K., Zhao, Q., Chen, H., Wang, J., and Xie, X · 2024
Later among the works it cites.
OWASP top ten 2025, 2025a
OWASP · 2025
Closest in time.
Sharma, M., Tong, M., Mu, J., Wei, J., Kruthoff, J., Goodfriend, S., Ong, E., Peng, A., Agarwal, R., Anil, C., et al · 2025
Closest in time.