2021

When Can Models Learn From Explanations? A Formal Framework for Understanding the Roles of Explanation Data

Hase, Peter, Bansal, Mohit

Understand

Many methods now exist for conditioning model outputs on task instructions, retrieved documents, and user-provided explanations and feedback.

  • Rather than relying solely on examples of task inputs and outputs, these approaches use valuable additional data for improving model correctness and aligning learned models with human priors.
  • Meanwhile, a growing body of evidence suggests that some language models can (1) store a large amount of knowledge in their parameters, and (2) perform inference over tasks in textual inputs at test time.
  • These results raise the possibility that, for some tasks, humans cannot explain to a model any more about the task than it already knows or could infer on its own.

Reading the bibliography…