Nachum, O., Chow, Y., Dai, B. & Li, L. (2019), Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections, in
2019
Later among the works it cites.
Duan, Y., Jia, Z. & Wang, M. (2020), Minimax-optimal off-policy evaluation with linear function approximation, in
2020
Later among the works it cites.
Feng, Y., Ren, T., Tang, Z. & Liu, Q. (2020), Accountable off-policy evaluation with kernel bellman statistics, in
2020
Later among the works it cites.
Kallus, N. & Uehara, M. (2020), Statistically efficient off-policy policy gradients, in
2020
Later among the works it cites.
Shi, C., Zhang, S., Lu, W. & Song, R. (2020), ‘Statistical inference of the value function for reinforcement learning in infinite horizon settings’, Journal of the Royal Statistical Society: Series B, in press
2020
Later among the works it cites.
Tang, Z., Feng, Y., Li, L., Zhou, D. & Liu, Q. (2020), Doubly robust bias reduction in infinite horizon off-policy estimation, in
2020
Later among the works it cites.
Uehara, M., Huang, J. & Jiang, N. (2020), Minimax weight and q-function learning for off-policy evaluation, in
2020
Later among the works it cites.
Zhang, S., Liu, B. & Whiteson, S. (2020), Gradientdice: Rethinking generalized offline estimation of stationary values, in
2020
Later among the works it cites.
Agarwal, A., Kakade, S. M., Lee, J. D. & Mahajan, G. (2021), ‘On the theory of policy gradient methods: Optimality, approximation, and distribution shift’, Journal of Machine Learning Research
2021
Later among the works it cites.
Chen, Y., Xu, L., Gulcehre, C., Paine, T. L., Gretton, A., de Freitas, N. & Doucet, A. (2021), ‘On instrumental variable regression for deep offline policy evaluation’, arXiv preprint arXiv:2105.10148
Original
2021
Later among the works it cites.
Duan, Y., Wang, M. & Wainwright, M. J. (2021), ‘Optimal policy evaluation using kernel-based temporal difference methods’, arXiv preprint arXiv:2109.12002
Original
2021
Later among the works it cites.
Jin, Y., Yang, Z. & Wang, Z. (2021), Is pessimism provably efficient for offline rl?, in
2021
Later among the works it cites.
Shi, C., Wan, R., Chernozhukov, V. & Song, R. (2021), Deeply-debiased off-policy interval estimation, in
2021
Later among the works it cites.
Uehara, M., Imaizumi, M., Jiang, N., Kallus, N., Sun, W. & Xie, T. (2021), ‘Finite sample analysis of minimax offline reinforcement learning: Completeness, fast rates and first-order efficiency’, arXiv preprint arXiv:2102.02981
Original
2021
Later among the works it cites.
Uehara, M. & Sun, W. (2021), ‘Pessimistic model-based offline rl: Pac bounds and posterior sampling under partial coverage’, arXiv preprint arXiv:2107.06226
Original
2021
Later among the works it cites.
Wang, J., Qi, Z. & Wong, R. K. (2021), ‘Projected state-action balancing weights for offline reinforcement learning’, arXiv preprint arXiv:2109.04640
Original
2021
Later among the works it cites.
Xie, T., Cheng, C.-A., Jiang, N., Mineiro, P. & Agarwal, A. (2021), ‘Bellman-consistent pessimism for offline reinforcement learning’, arXiv preprint arXiv:2106.06926
Original
2021
Later among the works it cites.
Xu, T., Yang, Z., Wang, Z. & Liang, Y. (2021), ‘Doubly robust off-policy actor-critic: Convergence and optimality’, arXiv preprint arXiv:2102.11866
Original
2021
Later among the works it cites.
Zanette, A., Wainwright, M. J. & Brunskill, E. (2021), ‘Provable benefits of actor-critic methods for offline reinforcement learning’, Advances in neural information processing systems
2021
Later among the works it cites.