Fetching the paper…
Reading the bibliography…
Infinite horizon off-policy policy evaluation is a highly challenging task due to the excessively large variance of typical importance sampling (IS) estimators.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Dynamic Programming and Optimal Control
Dimitri P. Bertsekas · 2000
Earlier work this paper cites.
Monte Carlo Strategies in Scientific Computing
Jun S. Liu · 2001
Earlier work this paper cites.
Marginal mean models for dynamic regimes
Susan A. Murphy, Mark van der Laan, and James M. Robins · 2001
Earlier work this paper cites.
The linear programming approach to approximate dynamic programming
Daniela Pucci de Farias and Benjamin Van Roy · 2003
Earlier work this paper cites.
Stochastic simulation: algorithms and analysis , volume 57
Søren Asmussen and Peter W Glynn · 2007
Earlier work this paper cites.
Learning from logged implicit exploration data
Alexander L. Strehl, John Langford, Lihong Li, and Sham M. Kakade · 2010
Earlier work this paper cites.
Doubly robust policy evaluation and learning
Miroslav Dudík, John Langford, and Lihong Li · 2011
Earlier work this paper cites.
Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms
Lihong Li, Wei Chu, John Langford, and Xuanhui Wang · 2011
Earlier work this paper cites.
Counterfactual reasoning and learning systems: The example of computational advertising
Léon Bottou, Jonas Peters, Joaquin Quiñonero-Candela, Denis Xavier Charles, D. Max Chickering, Elon Portugaly, Dipankar Ray, Patrice Simard, and Ed Snelson · 2013
Earlier work this paper cites.
Batch mode reinforcement learning based on the synthesis of artificial trajectories
Raphael Fonteneau, Susan A. Murphy, Louis Wehenkel, and Damien Ernst · 2013
Cited alongside, same era.
Markov Decision Processes.: Discrete Stochastic Dynamic Programming
Martin L Puterman · 2014
Cited alongside, same era.
Toward minimax off-policy value estimation
Lihong Li, Rémi Munos, and Csaba Szepesvári · 2015
Cited alongside, same era.
Finite-sample analysis of proximal gradient td algorithms
Bo Liu, Ji Liu, Mohammad Ghavamzadeh, Sridhar Mahadevan, and Marek Petrik · 2015
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Stochastic primal-dual methods and sample complexity of reinforcement learning
Consistent on-line off-policy evaluation
Assaf Hallak and Shie Mannor · 2017
Later among the works it cites.
Predictive off-policy policy evaluation for nonstationary decision problems, with applications to digital marketing
Philip S. Thomas, Georgios Theocharous, Mohammad Ghavamzadeh, Ishan Durugkar, and Emma Brunskill · 2017
Later among the works it cites.
Optimal and adaptive off-policy evaluation in contextual bandits
Yu-Xiang Wang, Alekh Agarwal, and Miroslav Dudík · 2017
Later among the works it cites.
More robust doubly robust off-policy evaluation
Mehrdad Farajtabar, Yinlam Chow, and Mohammad Ghavamzadeh · 2018
Later among the works it cites.
A kernel loss for solving the bellman equation
Yihao Feng, Lihong Li, and Qiang Liu · 2019
Closest in time.
Off-policy deep reinforcement learning by bootstrapping the covariate shift
Carles Gelada and Marc G Bellemare · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yichen Chen and Mengdi Wang · 2016
Cited alongside, same era.
Doubly robust off-policy evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2016
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
Philip S. Thomas and Emma Brunskill · 2016
Cited alongside, same era.
Using options and covariance testing for long horizon off-policy policy evaluation
Zhaohan Guo, Philip S. Thomas, and Emma Brunskill · 2017
Cited alongside, same era.
Learning from conditional distributions via dual kernel embeddings
Bo Dai, Niao He, Yunpeng Pan, Byron Boots, and Le Song
Cited in the paper.
Sbeed: Convergent reinforcement learning with nonlinear function approximation
Bo Dai, Albert Shaw, Lihong Li, Lin Xiao, Niao He, Zhen Liu, Jianshu Chen, and Le Song
Cited in the paper.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Qiang Liu, Lihong Li, Ziyang Tang, and Dengyong Zhou
Cited in the paper.
Closest in time.
Off-policy policy gradient with state distribution correction
Yao Liu, Adith Swaminathan, Alekh Agarwal, and Emma Brunskill · 2019
Closest in time.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
Ofir Nachum, Yinlam Chow, Bo Dai, and Lihong Li · 2019
Closest in time.
Optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
Tengyang Xie, Yifei Ma, and Yu-Xiang Wang · 2019
Closest in time.