2023

Challenges with unsupervised LLM knowledge discovery

Farquhar, Sebastian, Varma, Vikrant, Kenton, Zachary et al.

Understand

We show that existing unsupervised methods on large language model (LLM) activations do not discover knowledge -- instead they seem to discover whatever feature of the activations is most prominent.

  • The idea behind unsupervised knowledge elicitation is that knowledge satisfies a consistency structure, which can be used to discover knowledge.
  • We first prove theoretically that arbitrary features (not just knowledge) satisfy the consistency structure of a particular leading unsupervised knowledge-elicitation method, contrast-consistent search (Burns et al.
  • - arXiv:2212.03827).

Reading the bibliography…