Fetching the paper…

Optimistic Policy Optimization is Provably Efficient in Non-stationary MDPs · Around