2019

Categorizing Wireheading in Partially Embedded Agents

Majha, Arushi, Sarkar, Sayan, Zagami, Davide

Understand

$\textit{Embedded agents}$ are not explicitly separated from their environment, lacking clear I/O channels.

  • Such agents can reason about and modify their internal parts, which they are incentivized to shortcut or $\textit{wirehead}$ in order to achieve the maximal reward.
  • In this paper, we provide a taxonomy of ways by which wireheading can occur, followed by a definition of wirehead-vulnerable agents.
  • Starting from the fully dualistic universal agent AIXI, we introduce a spectrum of partially embedded agents and identify wireheading opportunities that such agents can exploit, experimentally demonstrating the results with the GRL simulation platform AIXIjs.

Reading the bibliography…