Fetching the paper…
Reading the bibliography…
We consider the problem of off-policy evaluation in Markov decision processes.
Monte carlo is fundamentally unsound
O’Hagan, A · 1987
Earlier work this paper cites.
Model-based direct adjustment
Rosenbaum, P. R · 1987
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, D., Sutton, R. S., and Singh, S · 2000
Earlier work this paper cites.
Bayesian monte carlo
Ghahramani, Z. and Rasmussen, C. E · 2003
Earlier work this paper cites.
Efficient estimation of average treatment effects using the estimated propensity score
Hirano, K., Imbens, G. W., and Ridder, G · 2003
Earlier work this paper cites.
Importance sampling via the estimated sampler
Henmi, M., Yoshida, R., and Eguchi, S · 2007
Earlier work this paper cites.
Doubly robust policy evaluation and learning
Dudík, M., Langford, J., and Li, L · 2011
Earlier work this paper cites.
Counterfactual reasoning and learning systems: The example of computational advertising
Bottou, L., Peters, J., Quiñonero-Candela, J., Charles, D. X., Chickering, D. M., Portugaly, E., Ray, D., Simard, P., and Snelson, E · 2013
Cited alongside, same era.
The cross-entropy method: a unified approach to combinatorial optimization, Monte-Carlo simulation and machine learning
Rubinstein, R. Y. and Kroese, D. P · 2013
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Cited alongside, same era.
Toward minimax off-policy value estimation
Li, L., Munos, R., and Szepesvári, C · 2015
Cited alongside, same era.
The self-normalized estimator for counterfactual learning
Swaminathan, A. and Joachims, T · 2015
Cited alongside, same era.
Safe Reinforcement Learning
Thomas, P. S · 2015
Openai baselines
Dhariwal, P., Hesse, C., Klimov, O., Nichol, A., Plappert, M., Radford, A., Schulman, J., Sidor, S., and Wu, Y · 2017
Later among the works it cites.
Importance sampling for fair policy selection
Doroudi, S., Thomas, P. S., and Brunskill, E · 2017
Later among the works it cites.
Data-efficient policy evaluation through behavior policy search
Hanna, J., Thomas, P. S., Stone, P., and Niekum, S · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Later among the works it cites.
More robust doubly robust off-policy evaluation
Farajtabar, M., Chow, Y., and Ghavamzadeh, M · 2018
Closest in time.
Reducing sampling error in the monte carlo policy gradient estimator
Hanna, J. and Stone, P · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Integral approximation by kernel smoothing
Delyon, B. and Portier, F · 2016
Cited alongside, same era.
Doubly robust off-policy evaluation for reinforcement learning
Jiang, N. and Li, L · 2016
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
Thomas, P. S. and Brunskill, E
Cited in the paper.
Magical policy search: Data efficient reinforcement learning with guarantees of global optimality
Thomas, P. S. and Brunskill, E
Cited in the paper.
Closest in time.
Efficient counterfactual learning from bandit feedback
Narita, Y., Yasui, S., and Yata, K · 2019
Closest in time.