2020

Offline Policy Selection under Uncertainty

Yang, Mengjiao, Dai, Bo, Nachum, Ofir et al.

Understand

The presence of uncertainty in policy evaluation significantly complicates the process of policy ranking and selection in real-world settings.

  • We formally consider offline policy selection as learning preferences over a set of policy prospects given a fixed experience dataset.
  • While one can select or rank policies based on point estimates of their policy values or high-confidence intervals, access to the full distribution over one's belief of the policy value enables more flexible selection algorithms under a wider range of downstream evaluation metrics.
  • We propose BayesDICE for estimating this belief distribution in terms of posteriors of distribution correction ratios derived from stochastic constraints (as opposed to explicit likelihood, which is not available).

Reading the bibliography…