2024

Extracting Prompts by Inverting LLM Outputs

Zhang, Collin, Morris, John X., Shmatikov, Vitaly

Understand

We consider the problem of language model inversion: given outputs of a language model, we seek to extract the prompt that generated these outputs.

  • We develop a new black-box method, output2prompt, that learns to extract prompts without access to the model's logits and without adversarial or jailbreaking queries.
  • In contrast to previous work, output2prompt only needs outputs of normal user queries.
  • To improve memory efficiency, output2prompt employs a new sparse encoding techique.

Reading the bibliography…