Understand
We consider the problem of language model inversion: given outputs of a language model, we seek to extract the prompt that generated these outputs.
- We develop a new black-box method, output2prompt, that learns to extract prompts without access to the model's logits and without adversarial or jailbreaking queries.
- In contrast to previous work, output2prompt only needs outputs of normal user queries.
- To improve memory efficiency, output2prompt employs a new sparse encoding techique.
Reading the bibliography…