2020

Minimax Value Interval for Off-Policy Evaluation and Policy Optimization

Jiang, Nan, Huang, Jiawei

Understand

We study minimax methods for off-policy evaluation (OPE) using value functions and marginalized importance weights.

  • Despite that they hold promises of overcoming the exponential variance in traditional importance sampling, several key problems remain: (1) They require function approximation and are generally biased.
  • For the sake of trustworthy OPE, is there anyway to quantify the biases? (2) They are split into two styles ("weight-learning" vs "value-learning").
  • Can we unify them? In this paper we answer both questions positively.

Reading the bibliography…