2017

Trial without Error: Towards Safe Reinforcement Learning via Human Intervention

Saunders, William, Sastry, Girish, Stuhlmueller, Andreas et al.

Understand

AI systems are increasingly applied to complex tasks that involve interaction with humans.

  • During training, such systems are potentially dangerous, as they haven't yet learned to avoid actions that could cause serious harm.
  • How can an AI system explore and learn without making a single mistake that harms humans or otherwise causes serious damage? For model-free reinforcement learning, having a human "in the loop" and ready to intervene is currently the only way to prevent all catastrophes.
  • We formalize human intervention for RL and show how to reduce the human labor required by training a supervised learner to imitate the human's intervention decisions.

Reading the bibliography…