2019

Efficiently Breaking the Curse of Horizon in Off-Policy Evaluation with Double Reinforcement Learning

Kallus, Nathan, Uehara, Masatoshi

Understand

Off-policy evaluation (OPE) in reinforcement learning is notoriously difficult in long- and infinite-horizon settings due to diminishing overlap between behavior and target policies.

  • In this paper, we study the role of Markovian and time-invariant structure in efficient OPE.
  • We first derive the efficiency bounds for OPE when one assumes each of these structures.
  • This precisely characterizes the curse of horizon: in time-variant processes, OPE is only feasible in the near-on-policy setting, where behavior and target policies are sufficiently similar.

Reading the bibliography…