2021

Human irrationality: both bad and good for reward inference

Chan, Lawrence, Critch, Andrew, Dragan, Anca

Understand

Assuming humans are (approximately) rational enables robots to infer reward functions by observing human behavior.

  • But people exhibit a wide array of irrationalities, and our goal with this work is to better understand the effect they can have on reward inference.
  • The challenge with studying this effect is that there are many types of irrationality, with varying degrees of mathematical formalization.
  • We thus operationalize irrationality in the language of MDPs, by altering the Bellman optimality equation, and use this framework to study how these alterations would affect inference.

Reading the bibliography…