2022

Models of human preference for learning reward functions

Knox, W. Bradley, Hatgis-Kessell, Stephane, Booth, Serena et al.

Understand

The utility of reinforcement learning is limited by the alignment of reward functions with the interests of human stakeholders.

  • One promising method for alignment is to learn the reward function from human-generated preferences between pairs of trajectory segments, a type of reinforcement learning from human feedback (RLHF).
  • These human preferences are typically assumed to be informed solely by partial return, the sum of rewards along each segment.
  • We find this assumption to be flawed and propose modeling human preferences instead as informed by each segment's regret, a measure of a segment's deviation from optimal decision-making.

Reading the bibliography…