Fetching the paper…
Reading the bibliography…
Motivated by the many real-world applications of reinforcement learning (RL) that require safe-policy iterations, we consider the problem of off-policy evaluation (OPE) -- the problem of evaluating a new policy using the historical data obtained by different behavior policies -- under the model of nonstationary episodic Markov Decision Processes (MDP) with a long horizon and a large action space.
A measure of asymptotic efficiency for tests of a hypothesis based on the sum of observations
Chernoff, H. et al. (1952) · 1952
Earlier work this paper cites.
The variance of discounted markov decision processes
Sobel, M. J. (1982) · 1982
Earlier work this paper cites.
Reinforcement learning with replacing eligibility traces
Singh, S. P. and Sutton, R. S. (1996) · 1996
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (1998) · 1998
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Precup, D., Sutton, R. S., and Singh, S. P. (2000) · 2000
Earlier work this paper cites.
Marginal mean models for dynamic regimes
Murphy, S. A., van der Laan, M. J., Robins, J. M., and Group, C. P. P. R. (2001) · 2001
Earlier work this paper cites.
Efficient estimation of average treatment effects using the estimated propensity score
Hirano, K., Imbens, G. W., and Ridder, G. (2003) · 2003
Earlier work this paper cites.
Clinical data based optimal sti strategies for hiv: a reinforcement learning approach
Ernst, D., Stan, G.-B., Goncalves, J., and Wehenkel, L. (2006) · 2006
Earlier work this paper cites.
Doubly robust policy evaluation and learning
Dudík, M., Langford, J., and Li, L. (2011) · 2011
Earlier work this paper cites.
Counterfactual reasoning and learning systems: The example of computational advertising
Bottou, L., Peters, J., Quiñonero-Candela, J., Charles, D. X., Chickering, D. M., Portugaly, E., Ray, D., Simard, P., and Snelson, E. (2013) · 2013
Cited alongside, same era.
Automatic ad format selection via contextual bandits
Tang, L., Rosales, R., Singh, A., and Agarwal, D. (2013) · 2013
Cited alongside, same era.
Offline policy evaluation across representations with applications to educational games
Mandel, T., Liu, Y.-E., Levine, S., Brunskill, E., and Popovic, Z. (2014) · 2014
Cited alongside, same era.
Simple and scalable response prediction for display advertising
Chapelle, O., Manavoglu, E., and Rosales, R. (2015) · 2015
Cited alongside, same era.
Toward minimax off-policy value estimation
Li, L., Munos, R., and Szepesvari, C. (2015) · 2015
Cited alongside, same era.
Personalized ad recommendation systems for life-time value optimization with guarantees
Data-efficient off-policy policy evaluation for reinforcement learning
Thomas, P. and Brunskill, E. (2016) · 2016
Later among the works it cites.
Using options and covariance testing for long horizon off-policy policy evaluation
Guo, Z., Thomas, P. S., and Brunskill, E. (2017) · 2017
Later among the works it cites.
Consistent on-line off-policy evaluation
Hallak, A. and Mannor, S. (2017) · 2017
Later among the works it cites.
Continuous state-space models for optimal sepsis treatment: a deep reinforcement learning approach
Raghu, A., Komorowski, M., Celi, L. A., Szolovits, P., and Ghassemi, M. (2017) · 2017
Later among the works it cites.
Predictive off-policy policy evaluation for nonstationary decision problems, with applications to digital marketing
Thomas, P. S., Theocharous, G., Ghavamzadeh, M., Durugkar, I., and Brunskill, E. (2017) · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Theocharous, G., Thomas, P. S., and Ghavamzadeh, M. (2015) · 2015
Cited alongside, same era.
Safe reinforcement learning
Thomas, P. S. (2015) · 2015
Cited alongside, same era.
High-confidence off-policy evaluation
Thomas, P. S., Theocharous, G., and Ghavamzadeh, M. (2015) · 2015
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Jiang, N. and Li, L. (2016) · 2016
Cited alongside, same era.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Liu, Q., Li, L., Tang, Z., and Zhou, D. (2018a)
Cited in the paper.
Representation balancing mdps for off-policy policy evaluation
Liu, Y., Gottesman, O., Raghu, A., Komorowski, M., Faisal, A. A., Doshi-Velez, F., and Brunskill, E. (2018b)
Cited in the paper.
Wang, Y.-X., Agarwal, A., and Dudık, M. (2017) · 2017
Later among the works it cites.
More robust doubly robust off-policy evaluation
Farajtabar, M., Chow, Y., and Ghavamzadeh, M. (2018) · 2018
Later among the works it cites.
Off-policy deep reinforcement learning by bootstrapping the covariate shift
Gelada, C. and Bellemare, M. G. (2019) · 2019
Closest in time.
Combining parametric and nonparametric models for off-policy evaluation
Gottesman, O., Liu, Y., Sussex, S., Brunskill, E., and Doshi-Velez, F. (2019) · 2019
Closest in time.