2024

Truth is Universal: Robust Detection of Lies in LLMs

Bürger, Lennart, Hamprecht, Fred A., Nadler, Boaz

Understand

Large Language Models (LLMs) have revolutionised natural language processing, exhibiting impressive human-like capabilities.

  • In particular, LLMs are capable of "lying", knowingly outputting false statements.
  • Hence, it is of interest and importance to develop methods to detect when LLMs lie.
  • Indeed, several authors trained classifiers to detect LLM lies based on their internal model activations.

Reading the bibliography…