Fetching the paper…
Reading the bibliography…
Off-policy estimation for long-horizon problems is important in many real-life applications such as healthcare and robotics, where high-fidelity simulators may not be available and on-policy evaluation is expensive or impossible.
RecSim: A configurable simulation platform for recommender systems, 2019
Eugene Ie, Chih-Wei Hsu, Martin Mladenov, Vihan Jain, Sanmit Narvekar, Jing Wang, Rui Wu, and Craig Boutilier · 1909
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Monte Carlo Strategies in Scientific Computing
Jun S. Liu · 2001
Earlier work this paper cites.
Marginal mean models for dynamic regimes
Susan A. Murphy, Mark van der Laan, and James M. Robins · 2001
Earlier work this paper cites.
Off-policy temporal-difference learning with funtion approximation
Doina Precup, Richard S. Sutton, and Sanjoy Dasgupta · 2001
Earlier work this paper cites.
Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method
Martin Riedmiller · 2005
Earlier work this paper cites.
Importance sampling via the estimated sampler
Masayuki Henmi, Ryo Yoshida, and Shinto Eguchi · 2007
Earlier work this paper cites.
Stable dual dynamic programming
Tao Wang, Daniel J. Lizotte, Michael H. Bowling, and Dale Schuurmans · 2007
Earlier work this paper cites.
Toward off-policy learning control with function approximation
Hamid Reza Maei, Csaba Szepesvári, Shalabh Bhatnagar, and Richard S. Sutton · 2010
Earlier work this paper cites.
Estimating divergence functionals and the likelihood ratio by convex risk minimization
XuanLong Nguyen, Martin J. Wainwright, and Michael I. Jordan · 2010
Earlier work this paper cites.
Reproducing kernel Hilbert spaces in probability and statistics
Alain Berlinet and Christine Thomas-Agnan · 2011
Earlier work this paper cites.
Doubly robust policy evaluation and learning
Miroslav Dudík, John Langford, and Lihong Li · 2011
Earlier work this paper cites.
A kernel two-sample test
Arthur Gretton, Karsten M Borgwardt, Malte J Rasch, Bernhard Schölkopf, and Alexander Smola · 2012
Earlier work this paper cites.
Counterfactual reasoning and learning systems: The example of computational advertising
Léon Bottou, Jonas Peters, Joaquin Quiñonero-Candela, Denis Xavier Charles, D. Max Chickering, Elon Portugaly, Dipankar Ray, Patrice Simard, and Ed Snelson · 2013
Earlier work this paper cites.
Batch mode reinforcement learning based on the synthesis of artificial trajectories
Raphael Fonteneau, Susan A. Murphy, Louis Wehenkel, and Damien Ernst · 2013
Cited alongside, same era.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Cited alongside, same era.
On using very large target vocabulary for neural machine translation
Sébastien Jean, KyungHyun Cho, Roland Memisevic, and Yoshua Bengio · 2015
Cited alongside, same era.
Toward minimax off-policy value estimation
Lihong Li, Remi Munos, and Csaba Szepesvári · 2015
Cited alongside, same era.
Doubly robust off-policy evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
Rémi Munos, Tom Stepleton, Anna Harutyunyan, and Marc G. Bellemare · 2016
Kernel mean embedding of distributions: A review and beyond
Krikamol Muandet, Kenji Fukumizu, Bharath Sriperumbudur, and Bernhard Schölkopf · 2017
Later among the works it cites.
Kernel mean embedding of distributions: A review and beyond
Krikamol Muandet, Kenji Fukumizu, Bharath Sriperumbudur, Bernhard Schölkopf, et al · 2017
Later among the works it cites.
Mengdi Wang · 2017
Later among the works it cites.
Boosting the actor with dual critic
Bo Dai, Albert Shaw, Niao He, Lihong Li, and Le Song · 2018
Later among the works it cites.
More robust doubly robust off-policy evaluation
Mehrdad Farajtabar, Yinlam Chow, and Mohammad Ghavamzadeh · 2018
Later among the works it cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
An emphatic approach to the problem of off-policy temporal-difference learning
Richard S. Sutton, A. Rupam Mahmood, and Martha White · 2016
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
Philip S. Thomas and Emma Brunskill · 2016
Cited alongside, same era.
Learning from conditional distributions via dual embeddings
Bo Dai, Niao He, Yunpeng Pan, Byron Boots, and Le Song · 2017
Cited alongside, same era.
Using options and covariance testing for long horizon off-policy policy evaluation
Zhaohan Guo, Philip S. Thomas, and Emma Brunskill · 2017
Cited alongside, same era.
Bootstrapping with models: Confidence intervals for off-policy evaluation
Josiah P. Hanna, Peter Stone, and Scott Niekum · 2017
Cited alongside, same era.
Markov Chains and Mixing Times
David A. Levin and Yuval Peres · 2017
Cited alongside, same era.
Qiang Liu, Lihong Li, Ziyang Tang, and Dengyong Zhou · 2018
Later among the works it cites.
Stein variational gradient descent as moment matching
Qiang Liu and Dilin Wang · 2018
Later among the works it cites.
Policy optimization via importance sampling
Alberto Maria Metelli, Matteo Papini, Francesco Faccio, and Marcello Restelli · 2018
Later among the works it cites.
Off-policy deep reinforcement learning by bootstrapping the covariate shift
Carles Gelada and Marc G. Bellemare · 2019
Later among the works it cites.
Off-policy policy gradient with state distribution correction
Yao Liu, Adith Swaminathan, Alekh Agarwal, and Emma Brunskill · 2019
Later among the works it cites.
DualDICE: Behavior-agnostic estimation of discounted stationary distribution corrections
Ofir Nachum, Yinlam Chow, Bo Dai, and Lihong Li · 2019
Later among the works it cites.
Optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
Tengyang Xie, Yifei Ma, and Yu-Xiang Wang · 2019
Later among the works it cites.
Generalized off-policy actor-critic
Shangtong Zhang, Wendelin Boehmer, and Shimon Whiteson · 2019
Later among the works it cites.