Fetching the paper…

Intrinsically Efficient, Stable, and Bounded Off-Policy Evaluation for Reinforcement Learning · Around