2023

Alignment for Honesty

Yang, Yuqing, Chern, Ethan, Qiu, Xipeng et al.

Understand

Recent research has made significant strides in aligning large language models (LLMs) with helpfulness and harmlessness.

  • In this paper, we argue for the importance of alignment for \emph{honesty}, ensuring that LLMs proactively refuse to answer questions when they lack knowledge, while still not being overly conservative.
  • However, a pivotal aspect of alignment for honesty involves discerning an LLM's knowledge boundaries, which demands comprehensive solutions in terms of metric development, benchmark creation, and training methodologies.
  • We address these challenges by first establishing a precise problem definition and defining ``honesty'' inspired by the Analects of Confucius.

Reading the bibliography…