2023

Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game

Toyer, Sam, Watkins, Olivia, Mendes, Ethan Adrian et al.

Understand

While Large Language Models (LLMs) are increasingly being used in real-world applications, they remain vulnerable to prompt injection attacks: malicious third party prompts that subvert the intent of the system designer.

  • To help researchers study this problem, we present a dataset of over 126,000 prompt injection attacks and 46,000 prompt-based "defenses" against prompt injection, all created by players of an online game called Tensor Trust.
  • To the best of our knowledge, this is currently the largest dataset of human-generated adversarial examples for instruction-following LLMs.
  • The attacks in our dataset have a lot of easily interpretable stucture, and shed light on the weaknesses of LLMs.

Reading the bibliography…