Fetching the paper…

Asymptotically Efficient Off-Policy Evaluation for Tabular Reinforcement Learning · Around