Fetching the paper…

Instance-optimality in optimal value estimation: Adaptivity via variance-reduced Q-learning · Around