The origins of phreaking, 2004
Gary D. Robson · 2004
Earlier work this paper cites.
The essence of command injection attacks in web applications
Zhendong Su and Gary Wassermann · 2006
Earlier work this paper cites.
A classification of SQL-injection attacks and countermeasures
William G Halfond, Jeremy Viegas, Alessandro Orso, et al · 2006
Earlier work this paper cites.
Cross site scripting prevention with dynamic data tainting and static analysis
Philipp Vogt, Florian Nentwich, Nenad Jovanovic, Engin Kirda, Christopher Kruegel, and Giovanni Vigna · 2007
Earlier work this paper cites.
On automated prepared statement generation to remove SQL injection vulnerabilities
Stephen Thomas, Laurie Williams, and Tao Xie · 2008
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2018
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le · 2021
Earlier work this paper cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Earlier work this paper cites.
Ignore previous prompt: Attack techniques for language models
Fábio Perez and Ian Ribeiro · 2022
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Original
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
Evaluating the susceptibility of pre-trained language models via handcrafted adversarial examples
Original
Hezekiah J Branch, Jonathan Rodriguez Cefalu, Jeremy McHugh, Leyla Hujer, Aditya Bahl, Daniel del Castillo Iglesias, Ron Heichman, and Ramesh Darwishi · 2022
Earlier work this paper cites.
Exploring prompt injection attacks, 2022
Jose Selvi · 2022
Earlier work this paper cites.
OpenAI Python API Library, 2022
OpenAI · 2022
Earlier work this paper cites.
GPT-4 Technical Report, 2023
OpenAI · 2023
Earlier work this paper cites.
Claude 2, 2023
Anthropic · 2023
Earlier work this paper cites.
Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection
Original
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz · 2023
Earlier work this paper cites.
OWASP Top 10 for LLM Applications, 2023
OWASP · 2023
Earlier work this paper cites.
Jatmo: Prompt injection defense by task-specific finetuning
Julien Piet, Maha Alrashed, Chawin Sitawarin, Sizhe Chen, Zeming Wei, Elizabeth Sun, Basel Alomair, and David Wagner · 2023
Earlier work this paper cites.
Jailbroken: How Does LLM Safety Training Fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Earlier work this paper cites.