Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have profoundly transformed natural language applications, with a growing reliance on instruction-based definitions for designing chatbots.
Amnesia: analysis and monitoring for neutralizing sql-injection attacks
William GJ Halfond and Alessandro Orso · 2005
Earlier work this paper cites.
A classification of sql-injection attacks and countermeasures
William G Halfond, Jeremy Viegas, Alessandro Orso, et al · 2006
Earlier work this paper cites.
The essence of command injection attacks in web applications
Zhendong Su and Gary Wassermann · 2006
Earlier work this paper cites.
Code injection attacks on harvard-architecture devices
Aurélien Francillon and Claude Castelluccia · 2008
Earlier work this paper cites.
Automated code injection prevention for web applications
Zhengqin Luo, Tamara Rezk, and Manuel Serrano · 2011
Earlier work this paper cites.
Predicting common web application vulnerabilities from input validation and sanitization code patterns
Lwin Khin Shar and Hee Beng Kuan Tan · 2012
Earlier work this paper cites.
Code injection attacks on html5-based mobile apps: Characterization, detection and mitigation
Xing Jin, Xunchao Hu, Kailiang Ying, Wenliang Du, Heng Yin, and Gautam Nagesh Peri · 2014
Earlier work this paper cites.
Precise client-side protection against { \{ DOM-based } \} { \{ Cross-Site } \} scripting
Ben Stock, Sebastian Lekies, Tobias Mueller, Patrick Spiegel, and Martin Johns · 2014
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Ernie: Enhanced representation through knowledge integration
Yu Sun, Shuohuan Wang, Yukun Li, Shikun Feng, Xuyi Chen, Han Zhang, Xin Tian, Danxiang Zhu, Hao Tian, and Hua Wu · 2019
Earlier work this paper cites.
Towards a human-like open-domain chatbot
Daniel Adiwardana, Minh-Thang Luong, David R So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, et al · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Earlier work this paper cites.
Gpt-neo: Large scale autoregressive language modeling with mesh-tensorflow
Sid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman · 2021
Earlier work this paper cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Earlier work this paper cites.
Demonstrate-search-predict: Composing retrieval and language models for knowledge-intensive NLP
Omar Khattab, Keshav Santhanam, Xiang Lisa Li, David Hall, Percy Liang, Christopher Potts, and Matei Zaharia · 2022
Earlier work this paper cites.
Blenderbot 3: a deployed conversational agent that continually learns to responsibly engage
Kurt Shuster, Jing Xu, Mojtaba Komeili, Da Ju, Eric Michael Smith, Stephen Roller, Megan Ung, Moya Chen, Kushal Arora, Joshua Lane, et al · 2022
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
Universal llm jailbreak: Chatgpt, gpt-4, bard, bing, anthropic, and beyond
Adversa AI · 2023
Cited alongside, same era.
Prompting is programming: A query language for large language models
Luca Beurer-Kellner, Marc Fischer, and Martin Vechev · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Kamar, et al · 2023
Cited alongside, same era.
Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz · 2023
Cited alongside, same era.
A list of leaked system prompts
Lakera AI (https://matt-rickard.com/a-list-of-leaked-system prompts) · 2023
Cited alongside, same era.
Niloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov, Maarten Sap, Reza Shokri, and Yejin Choi · 2023
Later among the works it cites.
Stealing the decoding algorithms of language models
Ali Naseh, Kalpesh Krishna, Mohit Iyyer, and Amir Houmansadr · 2023
Later among the works it cites.
Rodrigo Pedro, Daniel Castro, Paulo Carreira, and Nuno Santos · 2023
Later among the works it cites.
Maatphor: Automated variant analysis for prompt injection attacks
Ahmed Salem, Andrew Paverd, and Boris Köpf · 2023
Later among the works it cites.
Survey of vulnerabilities in large language models revealed by adversarial attacks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
OpenAI (https://openai.com/blog/introducing-the-gpt store) · 2023
Cited alongside, same era.
Claude-2
Anthropic AI (https://www.anthropic.com/news/claude 2) · 2023
Cited alongside, same era.
Gandalf ignore instructions
Lakera AI (https://www.lakera.ai) · 2023
Cited alongside, same era.
Langchain
(https://www.langchain.com/) · 2023
Cited alongside, same era.
Leaked prompt for microsoft bing
Twitter Post (http://tinyurl.com/3dy3mpm6) · 2023
Cited alongside, same era.
Leaked prompt for perplexity
Twitter Post (http://tinyurl.com/3wvcy484) · 2023
Cited alongside, same era.
Leaked prompt for myai from snap
Twitter Post (http://tinyurl.com/4mehurp9) · 2023
Cited alongside, same era.
Erfan Shayegani, Md Abdullah Al Mamun, Yu Fu, Pedram Zaree, Yue Dong, and Nael Abu-Ghazaleh · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson · 2023
Later among the works it cites.
Self-hardening firewall for large language models
Aegis (https://github.com/automorphic ai/aegis) · 2024
Closest in time.
Security scanner for large language model (llm) prompts
Vigil (https://github.com/deadbits/vigil llm) · 2024
Closest in time.
The security toolkit for llm interactions
LLMGuard (https://github.com/protectai/llm guard) · 2024
Closest in time.
Llm prompt injection detector
Rebuff (https://github.com/protectai/rebuff) · 2024
Closest in time.
automatically tests prompt injection attacks on chatgpt instances
Promptmap (https://github.com/utkusen/promptmap) · 2024
Closest in time.
An open-source toolkit for monitoring large language models
LangKit (https://github.com/whylabs/langkit) · 2024
Closest in time.
Collection of chatgpt jailbreak prompts
Jailbreak Chat (https://www.jailbreakchat.com/) · 2024
Closest in time.
Bringing enterprise-grade security to llms with one line of code
Lakera Guard (https://www.lakera.ai/blog/lakera-guard overview) · 2024
Closest in time.
Tensor Trust: Interpretable prompt injection attacks from an online game
Sam Toyer, Olivia Watkins, Ethan Adrian Mendes, Justin Svegliato, Luke Bailey, Tiffany Wang, Isaac Ong, Karim Elmaaroufi, Pieter Abbeel, Trevor Darrell, Alan Ritter, and Stuart Russell · 2024
Closest in time.