2020

REALab: An Embedded Perspective on Tampering

Kumar, Ramana, Uesato, Jonathan, Ngo, Richard et al.

Understand

This paper describes REALab, a platform for embedded agency research in reinforcement learning (RL).

  • REALab is designed to model the structure of tampering problems that may arise in real-world deployments of RL.
  • Standard Markov Decision Process (MDP) formulations of RL and simulated environments mirroring the MDP structure assume secure access to feedback (e.g., rewards).
  • This may be unrealistic in settings where agents are embedded and can corrupt the processes producing feedback (e.g., human supervisors, or an implemented reward function).

Reading the bibliography…