Fetching the paper…
Reading the bibliography…
We study the problem of off-policy value evaluation in reinforcement learning (RL), where one aims to estimate the value of a new policy based on data collected by a different policy.
Statistics and causal inference
Holland, Paul W · 1986
Earlier work this paper cites.
Semiparametric regression estimation in the presence of dependent censoring
Rotnitzky, Andrea and Robins, James M · 1995
Earlier work this paper cites.
Neuro-Dynamic Programming (Optimization and Neural Computation Series, 3)
Bertsekas, Dimitri P and Tsitsiklis, John N · 1996
Earlier work this paper cites.
Reinforcement learning with replacing eligibility traces
Singh, Satinder P and Sutton, Richard S · 1996
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, Richard S. and Barto, Andrew G · 1998
Earlier work this paper cites.
The UCI KDD Archive
Hettich, S. and Bay, S. D · 1999
Earlier work this paper cites.
Temporal abstraction in reinforcement learning
Precup, Doina · 2000
Earlier work this paper cites.
Eligibility Traces for Off-Policy Policy Evaluation
Precup, Doina, Sutton, Richard S., and Singh, Satinder P · 2000
Earlier work this paper cites.
Marginal Mean Models for Dynamic Regimes
Murphy, Susan A., van der Laan, Mark, and Robins, James M · 2001
Earlier work this paper cites.
Off-Policy Temporal-Difference Learning with Funtion Approximation
Precup, Doina, Sutton, Richard S., and Dasgupta, Sanjoy · 2001
Earlier work this paper cites.
Approximately Optimal Approximate Reinforcement Learning
Kakade, Sham and Langford, John · 2002
Earlier work this paper cites.
Kernel-based reinforcement learning
Ormoneit, Dirk and Sen, Śaunak · 2002
Cited alongside, same era.
Efficient estimation of average treatment effects using the estimated propensity score
Hirano, Keisuke, Imbens, Guido W., and Ridder, Geert · 2003
Cited alongside, same era.
Bandit based monte-carlo planning
Kocsis, Levente and Szepesvári, Csaba · 2006
Cited alongside, same era.
Model-based function approximation in reinforcement learning
Jong, Nicholas K and Stone, Peter · 2007
Cited alongside, same era.
Bias and variance approximation in value function estimates
Mannor, Shie, Simester, Duncan, Sun, Peng, and Tsitsiklis, John N · 2007
Cited alongside, same era.
Causality: Models, Reasoning, and Inference
Pearl, Judea · 2009
Cited alongside, same era.
Counterfactual Reasoning and Learning Systems: The Example of Computational Advertising
Bottou, Léon, Peters, Jonas, Quiñonero-Candela, Joaquin, Charles, Denis Xavier, Chickering, D. Max, Portugaly, Elon, Ray, Dipankar, Simard, Patrice, and Snelson, Ed · 2013
Later among the works it cites.
Batch mode reinforcement learning based on the synthesis of artificial trajectories
Fonteneau, Raphael, Murphy, Susan A., Wehenkel, Louis, and Ernst, Damien · 2013
Later among the works it cites.
Off-policy Evaluation in Markov Decision Processes
Paduraru, Cosmin · 2013
Later among the works it cites.
Safe policy iteration
Pirotta, Matteo, Restelli, Marcello, Pecorino, Alessio, and Calandriello, Daniele · 2013
Later among the works it cites.
Policy evaluation with temporal differences: A survey and comparison
Dann, Christoph, Neumann, Gerhard, and Peters, Jan · 2014
Later among the works it cites.
Offline policy evaluation across representations with applications to educational games
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A theory of Cramer-Rao bounds for constrained parametric models
Moore Jr, Terrence Joseph · 2010
Cited alongside, same era.
Doubly Robust Policy Evaluation and Learning
Dudík, Miroslav, Langford, John, and Li, Lihong · 2011
Cited alongside, same era.
Model selection in reinforcement learning
Farahmand, Amir-massoud and Szepesvári, Csaba · 2011
Cited alongside, same era.
Unbiased Offline Evaluation of Contextual-bandit-based News Article Recommendation Algorithms
Li, Lihong, Chu, Wei, Langford, John, and Wang, Xuanhui · 2011
Cited alongside, same era.
Modelling transition dynamics in MDPs with RKHS embeddings
Grünewälder, S, Lever, G, Baldassarre, L, Pontil, M, and Gretton, A · 2012
Cited alongside, same era.
Toward minimax off-policy value estimation
Li, Lihong, Munos, Remi, and Szepesvári, Csaba
Cited in the paper.
Mandel, Travis, Liu, Yun-En, Levine, Sergey, Brunskill, Emma, and Popovic, Zoran · 2014
Later among the works it cites.
Abstraction Selection in Model-based Reinforcement Learning
Jiang, Nan, Kulesza, Alex, and Singh, Satinder · 2015
Closest in time.
Improved Empirical Methods in Reinforcement-Learning Evaluation
Marivate, Vukosi N · 2015
Closest in time.
An emphatic approach to the problem of off-policy temporal-difference learning
Sutton, Richard S., Mahmood, Ashique Rupam, and White, Martha · 2015
Closest in time.
Safe Reinforcement Learning
Thomas, Philip · 2015
Closest in time.
Data-Efficient Off-Policy Policy Evaluation for Reinforcement Learning
Thomas, Philip and Brunskill, Emma · 2016
Closest in time.