Fetching the paper…
Reading the bibliography…
We consider the problem of off-policy evaluation for reinforcement learning, where the goal is to estimate the expected reward of a target policy $\pi$ using offline data collected by running a logging policy $\mu$.
Combining parametric and nonparametric models for off-policy evaluation
Gottesman, O., Liu, Y., Sussex, S., Brunskill, E., & Doshi-Velez, F. (2019) · 1905
Earlier work this paper cites.
On the optimality of sparse model-based planning for markov decision processes
Agarwal, A., Kakade, S., & Yang, L. F. (2019) · 1906
Earlier work this paper cites.
Double reinforcement learning for efficient off-policy evaluation in markov decision processes
Kallus, N., & Uehara, M. (2019a) · 1908
Earlier work this paper cites.
Kallus, N., & Uehara, M. (2019b) · 1909
Earlier work this paper cites.
A measure of asymptotic efficiency for tests of a hypothesis based on the sum of observations
Chernoff, H., et al. (1952) · 1952
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
Puterman, M. L. (1994) · 1994
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S., & Barto, A. G. (1998) · 1998
Earlier work this paper cites.
Finite-sample convergence rates for q-learning and indirect algorithms
Kearns, M. J., & Singh, S. P. (1999) · 1999
Earlier work this paper cites.
Asymptotic statistics
Van der Vaart, A. W. (2000) · 2000
Earlier work this paper cites.
Marginal mean models for dynamic regimes
Murphy, S. A., van der Laan, M. J., Robins, J. M., & Group, C. P. P. R. (2001) · 2001
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, M., & Singh, S. (2002) · 2002
Earlier work this paper cites.
A gentle introduction to concentration inequalities
Sridharan, K. (2002) · 2002
Earlier work this paper cites.
Efficient estimation of average treatment effects using the estimated propensity score
Hirano, K., Imbens, G. W., & Ridder, G. (2003) · 2003
Cited alongside, same era.
Doubly robust policy evaluation and learning
Dudík, M., Langford, J., & Li, L. (2011) · 2011
Cited alongside, same era.
Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms
Li, L., Chu, W., Langford, J., & Wang, X. (2011) · 2011
Cited alongside, same era.
Minimax pac bounds on the sample complexity of reinforcement learning with a generative model
Azar, M. G., Munos, R., & Kappen, H. J. (2013) · 2013
Cited alongside, same era.
Weighted importance sampling for off-policy learning with linear function approximation
Mahmood, A. R., van Hasselt, H. P., & Sutton, R. S. (2014) · 2014
Cited alongside, same era.
Consistent on-line off-policy evaluation
Hallak, A., & Mannor, S. (2017) · 2017
Later among the works it cites.
Off-policy evaluation for slate recommendation
Swaminathan, A., Krishnamurthy, A., Agarwal, A., Dudik, M., Langford, J., Jose, D., & Zitouni, I. (2017) · 2017
Later among the works it cites.
Predictive off-policy policy evaluation for nonstationary decision problems, with applications to digital marketing
Thomas, P. S., Theocharous, G., Ghavamzadeh, M., Durugkar, I., & Brunskill, E. (2017) · 2017
Later among the works it cites.
Optimal and adaptive off-policy evaluation in contextual bandits
Wang, Y.-X., Agarwal, A., & Dudik, M. (2017) · 2017
Later among the works it cites.
More robust doubly robust off-policy evaluation
Farajtabar, M., Chow, Y., & Ghavamzadeh, M. (2018) · 2018
Later among the works it cites.
Notes on tabular methods
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Offline policy evaluation across representations with applications to educational games
Mandel, T., Liu, Y.-E., Levine, S., Brunskill, E., & Popovic, Z. (2014) · 2014
Cited alongside, same era.
Toward minimax off-policy value estimation
Li, L., Munos, R., & Szepesvári, C. (2015) · 2015
Cited alongside, same era.
Personalized ad recommendation systems for life-time value optimization with guarantees
Theocharous, G., Thomas, P. S., & Ghavamzadeh, M. (2015) · 2015
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Jiang, N., & Li, L. (2016) · 2016
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
Thomas, P., & Brunskill, E. (2016) · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., & Munos, R. (2017) · 2017
Cited alongside, same era.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Dann, C., Lattimore, T., & Brunskill, E. (2017) · 2017
Cited alongside, same era.
Jiang, N. (2018) · 2018
Later among the works it cites.
Open problem: The dependence of sample complexity lower bounds on planning horizon
Jiang, N., & Agarwal, A. (2018) · 2018
Later among the works it cites.
Is q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., & Jordan, M. I. (2018) · 2018
Later among the works it cites.
Near-optimal time and sample complexities for solving markov decision processes with a generative model
Sidford, A., Wang, M., Wu, X., Yang, L., & Ye, Y. (2018) · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S., & Barto, A. G. (2018) · 2018
Later among the works it cites.
Off-policy deep reinforcement learning by bootstrapping the covariate shift
Gelada, C., & Bellemare, M. G. (2019) · 2019
Later among the works it cites.
Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
Xie, T., Ma, Y., & Wang, Y.-X. (2019) · 2019
Later among the works it cites.