Fetching the paper…
Reading the bibliography…
We consider the off-policy estimation problem of estimating the expected reward of a target policy using samples collected by a different behavior policy.
Dynamic Programming
Richard E. Bellman · 1957
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Hierarchical reinforcement learning with the MAXQ value function decomposition
Thomas G Dietterich · 2000
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup, Richard S. Sutton, and Satinder P. Singh · 2000
Earlier work this paper cites.
Monte Carlo Strategies in Scientific Computing
Jun S. Liu · 2001
Earlier work this paper cites.
Marginal mean models for dynamic regimes
Susan A. Murphy, Mark van der Laan, and James M. Robins · 2001
Earlier work this paper cites.
Off-policy temporal-difference learning with funtion approximation
Doina Precup, Richard S. Sutton, and Sanjoy Dasgupta · 2001
Earlier work this paper cites.
Learning with kernels: support vector machines, regularization, optimization, and beyond
Bernhard Scholkopf and Alexander J Smola · 2001
Earlier work this paper cites.
Efficient estimation of average treatment effects using the estimated propensity score
Keisuke Hirano, Guido W Imbens, and Geert Ridder · 2003
Earlier work this paper cites.
Least-squares policy iteration
Michail G. Lagoudakis and Ronald Parr · 2003
Earlier work this paper cites.
Stochastic Simulation: Algorithms and Analysis
Søren Asmussen and Peter W. Glynn · 2007
Earlier work this paper cites.
Approximation theorems of mathematical statistics
Robert J Serfling · 2009
Earlier work this paper cites.
Estimating divergence functionals and the likelihood ratio by convex risk minimization
XuanLong Nguyen, Martin J Wainwright, and Michael Jordan · 2010
Earlier work this paper cites.
Learning from logged implicit exploration data
Alexander L. Strehl, John Langford, Lihong Li, and Sham M. Kakade · 2010
Cited alongside, same era.
Reproducing kernel Hilbert spaces in probability and statistics
Alain Berlinet and Christine Thomas-Agnan · 2011
Cited alongside, same era.
Doubly robust policy evaluation and learning
Miroslav Dudík, John Langford, and Lihong Li · 2011
Cited alongside, same era.
Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms
Lihong Li, Wei Chu, John Langford, and Xuanhui Wang · 2011
Cited alongside, same era.
A kernel two-sample test
Arthur Gretton, Karsten M Borgwardt, Malte J Rasch, Bernhard Schölkopf, and Alexander Smola · 2012
Cited alongside, same era.
Recent development and applications of sumo-simulation of urban mobility
Daniel Krajzewicz, Jakob Erdmann, Michael Behrisch, and Laura Bieker · 2012
Toward minimax off-policy value estimation
Lihong Li, Rémi Munos, and Csaba Szepesvári · 2015
Later among the works it cites.
Generalized emphatic temporal difference learning: Bias-variance analysis
Assaf Hallak, Aviv Tamar, Remi Munos, and Shie Mannor · 2016
Later among the works it cites.
Doubly robust off-policy evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2016
Later among the works it cites.
Safe and efficient off-policy reinforcement learning
Rémi Munos, Tom Stepleton, Anna Harutyunyan, and Marc G. Bellemare · 2016
Later among the works it cites.
An emphatic approach to the problem of off-policy temporal-difference learning
Richard S. Sutton, A. Rupam Mahmood, and Martha White · 2016
Later among the works it cites.
Data-efficient off-policy policy evaluation for reinforcement learning
Philip S. Thomas and Emma Brunskill · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Density ratio estimation in machine learning
Masashi Sugiyama, Taiji Suzuki, and Takafumi Kanamori · 2012
Cited alongside, same era.
Counterfactual reasoning and learning systems: The example of computational advertising
Léon Bottou, Jonas Peters, Joaquin Quiñonero-Candela, Denis Xavier Charles, D. Max Chickering, Elon Portugaly, Dipankar Ray, Patrice Simard, and Ed Snelson · 2013
Cited alongside, same era.
Monte Carlo Theory, Methods and Examples
Art B. Owen · 2013
Cited alongside, same era.
Automatic ad format selection via contextual bandits
Liang Tang, Romer Rosales, Ajit Singh, and Deepak Agarwal · 2013
Cited alongside, same era.
Simple and scalable response prediction for display advertising
Olivier Chapelle, Eren Manavoglu, and Romer Rosales · 2014
Cited alongside, same era.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Cited alongside, same era.
Later among the works it cites.
Coordinated deep reinforcement learners for traffic light control
Elise Van der Pol and Frans A Oliehoek · 2016
Later among the works it cites.
Using options and covariance testing for long horizon off-policy policy evaluation
Zhaohan Guo, Philip S. Thomas, and Emma Brunskill · 2017
Later among the works it cites.
Consistent on-line off-policy evaluation
Assaf Hallak and Shie Mannor · 2017
Later among the works it cites.
Markov chains and mixing times
David A Levin and Yuval Peres · 2017
Later among the works it cites.
Kernel mean embedding of distributions: A review and beyond
Krikamol Muandet, Kenji Fukumizu, Bharath Sriperumbudur, Bernhard Schölkopf, et al · 2017
Later among the works it cites.
Predictive off-policy policy evaluation for nonstationary decision problems, with applications to digital marketing
Philip S. Thomas, Georgios Theocharous, Mohammad Ghavamzadeh, Ishan Durugkar, and Emma Brunskill · 2017
Later among the works it cites.
Optimal and adaptive off-policy evaluation in contextual bandits
Yu-Xiang Wang, Alekh Agarwal, and Miroslav Dudík · 2017
Later among the works it cites.
Action-dependent control variates for policy optimization via stein identity
Hao Liu, Yihao Feng, Yi Mao, Dengyong Zhou, Jian Peng, and Qiang Liu · 2018
Closest in time.