Fetching the paper…

A Minimax Learning Approach to Off-Policy Evaluation in Confounded Partially Observable Markov Decision Processes · Around