Fetching the paper…
Reading the bibliography…
We surface a new threat to closed-weight Large Language Models (LLMs) that enables an attacker to compute optimization-based prompt injections.
F. Tramèr, F. Zhang, A. Juels, M. K. Reiter, and T. Ristenpart, “Stealing machine learning models via prediction { \{ APIs } \} ,” in 25th USENIX security symposium (USENIX Security 16) , 2016, pp. 601–618
2016
Earlier work this paper cites.
J. Wei, Y. Zhang, Z. Zhou, Z. Li, and M. A. Al Faruque, “Leaky dnn: Stealing deep-learning model secret with gpu context-switching side-channel,” in 2020 50th Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN) . IEEE, 2020, pp. 125–137
2020
Earlier work this paper cites.
L. Gao, “On the sizes of openai api models,” https://blog.eleuther.ai/gpt3-model-sizes/ , 2021, accessed: [Date Accessed]
2021
Earlier work this paper cites.
J. Wei, M. P. Bosma, V. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le, “Finetuned language models are zero-shot learners,” 2022. [Online]. Available: https://openreview.net/forum?id=gEZrGCozdqR
2022
Earlier work this paper cites.
F. Perez and I. Ribeiro, “Ignore previous prompt: Attack techniques for language models,” 2022
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
N. Jain, A. Schwarzschild, Y. Wen, G. Somepalli, J. Kirchenbauer, P. yeh Chiang, M. Goldblum, A. Saha, J. Geiping, and T. Goldstein, “Baseline defenses for adversarial attacks against aligned language models,” 2023
2023
Earlier work this paper cites.
A. Wei, N. Haghtalab, and J. Steinhardt, “Jailbroken: How does llm safety training fail?” 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, “Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection,” 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
R. Samoilenko, “New prompt injection attack on chatgpt web version. markdown images can steal your chat data.” 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
A. Mehrotra, M. Zampetakis, P. Kassianik, B. Nelson, H. Anderson, Y. Singer, and A. Karbasi, “Tree of attacks: Jailbreaking black-box llms automatically,” 2023
2023
Earlier work this paper cites.
A. Zou, Z. Wang, J. Z. Kolter, and M. Fredrikson, “Universal and transferable adversarial attacks on aligned language models,” 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
J. Rehberger, “Ai injections: Direct and indirect prompt injections and their implications,” https://embracethered.com/blog/posts/2023/ai-injections-direct-and-indirect-prompt-injection-basics/ , 2023
2023
Earlier work this paper cites.
S. Willison, “Prompt injection: What’s the worst that can happen?” https://simonwillison.net/2023/Apr/14/worst-that-can-happen/ , 2023
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
A. Wan, E. Wallace, S. Shen, and D. Klein, “Poisoning language models during instruction tuning,” in International Conference on Machine Learning . PMLR, 2023, pp. 35 413–35 425
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Google, “Generating content — Gemini API,” https://ai.google.dev/api/generate-content#generatecontentresponse , 2024, [Accessed 23-09-2024]
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
SnykSec, “Agent hijacking: The true impact of prompt injection attacks,” https://dev.to/snyk/agent-hijacking-the-true-impact-of-prompt-injection-attacks-983 , 2024, [Accessed 23-09-2024]
2024
Cited alongside, same era.
Google, “Fine-tuning with the Gemini API — Google AI for Developers — ai.google.dev,” https://ai.google.dev/gemini-api/docs/model-tuning , 2024, [Accessed 23-09-2024]
2024
Cited alongside, same era.
OpenAI, “Fine-tuning now available for gpt-4o,” https://openai.com/index/gpt-4o-fine-tuning/ , 2024, [Accessed 22-09-2024]
2024
Cited alongside, same era.
Amazon Web Services (AWS), “Fine-tune anthropic’s claude 3 haiku in amazon bedrock to boost model accuracy and quality,” https://aws.amazon.com/blogs/machine-learning/fine-tune-anthropics-claude-3-haiku-in-amazon-bedrock-to-boost-model-accuracy-and-quality/ , 2023, accessed: 2024-11-14
2024
Cited alongside, same era.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
P. Sebastian Raschka, “Llm research insights: Instruction masking and new lora finetuning experiments,” https://www.linkedin.com/pulse/llm-research-insights-instruction-masking-new-lora-raschka-phd-7p1oc/ , Jun. 2024, accessed: 2024-11-14
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
G. C. AI, “Model tuning with gemini api,” https://ai.google.dev/gemini-api/docs/model-tuning , 2023, accessed: 2024-11-14
2024
Later among the works it cites.
Anyscale, “Fine-tuning llms: Lora or full parameter? an in-depth analysis with llama 2,” https://www.anyscale.com/blog/fine-tuning-llms-lora-or-full-parameter-an-in-depth-analysis-with-llama-2 , 2023, accessed: 2024-11-14
2024
Later among the works it cites.
C. A. et al., “Many-shot jailbreaking — anthropic.com,” https://www.anthropic.com/research/many-shot-jailbreaking , 2024, [Accessed 27-09-2024]
2024
Later among the works it cites.
K. Robison, “OpenAI’s latest model will block the ‘ignore all previous instructions’ loophole,” https://www.theverge.com/2024/7/19/24201414/openai-chatgpt-gpt-4o-prompt-injection-instruction-hierarchy , 2024, [Accessed 27-09-2024]
2024
Later among the works it cites.
2024
Later among the works it cites.
Agentic AI Security Team at Google DeepMind, “How we estimate the risk from prompt injection attacks on ai systems,” https://security.googleblog.com/2025/01/how-we-estimate-risk-from-prompt.html , Jan. 2025, [Accessed 29-01-2025]
2025
Closest in time.
Anthropic, “Fine-tune claude 3 haiku,” https://www.anthropic.com/news/fine-tune-claude-3-haiku , 2024, [Accessed 31-03-2025]
2025
Closest in time.