Brockman G, Cheung V, Pettersson L, Schneider J, Schulman J, Tang J, Zaremba W (2016) OpenAI gym. arXiv preprint arXiv:1606.01540
Original
2016
Later among the works it cites.
Jiang N, Li L (2016) Doubly robust off-policy value evaluation for reinforcement learning. In Proceedings of the 33rd International Conference on International Conference on Machine Learning-Volume 652–661
2016
Later among the works it cites.
Munos R, Stepleton T, Harutyunyan A, Bellemare M (2016) Safe and efficient off-policy reinforcement learning. Advances in Neural Information Processing Systems 29 , 1054–1062
2016
Later among the works it cites.
Thomas P, Brunskill E (2016) Data-efficient off-policy policy evaluation for reinforcement learning. In Proceedings of the 33rd International Conference on Machine Learning 2139–2148
2016
Later among the works it cites.
Chernozhukov V, Chetverikov D, Demirer M, Duflo E, Hansen C, Newey W, Robins J (2018) Double/debiased machine learning for treatment and structural parameters. Econometrics Journal 21:C1–C68
2018
Later among the works it cites.
Farajtabar M, Chow Y, Ghavamzadeh M (2018) More robust doubly robust off-policy evaluation. In Proceedings of the 35th International Conference on Machine Learning 1447–1456
2018
Later among the works it cites.
Kallus N (2018) Balanced policy evaluation and learning. Advances in Neural Information Processing Systems , 8895–8906
2018
Later among the works it cites.
Luckett DJ, Laber EB, Kahkoska AR, Maahs DM, Mayer-Davis E, Kosorok MR (2018) Estimating dynamic treatment regimes in mobile health using v-learning. Journal of the American Statistical Association 1–34
2018
Later among the works it cites.
Sutton RS, Barto AG (2018) Reinforcement learning: An introduction (Cambridge: MIT press)
2018
Later among the works it cites.
Gottesman O, Johansson F, Komorowski M, Faisal A, Sontag D, Doshi-Velez F, Celi LA (2019) Guidelines for reinforcement learning in healthcare. Nat Med 25:16–18
2019
Closest in time.
Kallus N, Uehara M (2019) Intrinsically efficient, stable, and bounded off-policy evaluation for reinforcement learning. Advances in Neural Information Processing Systems 32 , 3320–3329
2019
Closest in time.
Nachum O, Chow Y, Dai B, Li L (2019) Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections. Advances in Neural Information Processing Systems 2019 (To appear)
2019
Closest in time.
Rotnitzky A, Smucler E, Robins J (2019) Characterization of parameters with a mixed bias property. arXiv preprint arXiv:1509.02556
Original
2019
Closest in time.
Xie T, Ma Y, Wang YX (2019) Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling. Advances in Neural Information Processing Systems 32 , 9665–9675
2019
Closest in time.
Huang J, Jiang N (2020) From importance sampling to doubly robust policy gradient. International Conference on Machine Learning , 4434–4443 (PMLR)
2020
Closest in time.
Tang Z, Feng Y, Li L, Zhou D, Liu Q (2020) Harnessing infinite-horizon off-policy evaluation: Double robustness via duality. ICLR 2020 (To appear)
2020
Closest in time.
Uehara M, Huang J, Jiang N (2020) Minimax weight and q-function learning for off-policy evaluation. International Conference on Machine Learning , 9659–9668
2020
Closest in time.
Ueno T, Kawanabe M, Mori T, Maeda SI, Ishii S (2011) Generalized td learning. Journal of Machine Learning Research 12:1977––2020
2020
Closest in time.
Huang A, Jiang N (2022) Beyond the return: Off-policy function estimation under user-specified error-measuring distributions. arXiv preprint arXiv:2210.15543
Original
2022
Closest in time.
Khan S, Tamer E (2010) Irregular identification, support conditions, and inverse weight estimation. Econometrica 78:2021–2042
2042
Closest in time.