2022

Uncertainty Estimation for Language Reward Models

Gleave, Adam, Irving, Geoffrey

Understand

Language models can learn a range of capabilities from unsupervised training on text corpora.

  • However, to solve a particular problem (such as text summarization) it is typically necessary to fine-tune them on a task-specific dataset.
  • It is often easier for humans to choose between options than to provide labeled data, and prior work has achieved state-of-the-art performance by training a reward model from such preference comparisons.
  • However, collecting a large preference comparison dataset is still expensive -- and the learned reward models are unreliable out-of-distribution.

Reading the bibliography…