Fetching the paper…
Reading the bibliography…
Off-policy policy estimators that use importance sampling (IS) can suffer from high variance in long-horizon domains, and there has been particular excitement over new IS methods that leverage the structure of Markov decision processes.
Double reinforcement learning for efficient off-policy evaluation in markov decision processes
Kallus, N. and Uehara, M · 1908
Earlier work this paper cites.
Kallus, N. and Uehara, M · 1909
Earlier work this paper cites.
Conditional monte carlo
Hammersley, J. M · 1956
Earlier work this paper cites.
The strong law of large numbers for a class of markov chains
Breiman, L · 1960
Earlier work this paper cites.
Weighted sums of certain dependent random variables
Azuma, K · 1967
Earlier work this paper cites.
The interpretation of conditional monte carlo as a form of importance sampling
Dubi, A. and Horowitz, Y. S · 1979
Earlier work this paper cites.
Optimal formulae of the conditional monte carlo
Granovsky, B. L · 1981
Earlier work this paper cites.
Simulation and the Monte Carlo Method
Rubinstein, R. Y · 1981
Earlier work this paper cites.
A converse to scheffe’s theorem
Boos, D. D. et al · 1985
Earlier work this paper cites.
A Guide to Simulation (2Nd Ed.)
Bratley, P., Fox, B. L., and Schrage, L. E · 1987
Earlier work this paper cites.
Simulation methods for queues: An overview
Glynn, P. W. and Iglehart, D. L · 1988
Earlier work this paper cites.
Advances in Importance Sampling
Hesterberg, T. C · 1988
Earlier work this paper cites.
Simulating average delay–variance reduction by conditioning
Ross, S. M · 1988
Earlier work this paper cites.
Filtered monte carlo
Glasserman, P · 1993
Cited alongside, same era.
Importance sampling for markov chains: asymptotics for the variance
Glynn, P. W · 1994
Cited alongside, same era.
Efficiency improvement and variance reduction
L’Ecuyer, P · 1994
Cited alongside, same era.
A liapounov bound for solutions of the poisson equation
Glynn, P. W., Meyn, S. P., et al · 1996
Cited alongside, same era.
Some results in importance sampling and an application to detection
Srinivasan, R · 1998
Cited alongside, same era.
Eligibility traces for off-policy policy evaluation
Precup, D., Sutton, R. S., and Singh, S. P · 2000
Cited alongside, same era.
Learning from scarce experience
Doubly robust off-policy value evaluation for reinforcement learning
Jiang, N. and Li, L · 2015
Later among the works it cites.
Double machine learning for treatment and causal parameters
Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., and Newey, W. K · 2016
Later among the works it cites.
Using options and covariance testing for long horizon off-policy policy evaluation
Guo, Z., Thomas, P. S., and Brunskill, E · 2017
Later among the works it cites.
Consistent on-line off-policy evaluation
Hallak, A. and Mannor, S · 2017
Later among the works it cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Liu, Q., Li, L., Tang, Z., and Zhou, D · 2018
Later among the works it cites.
Policy optimization via importance sampling
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Peshkin, L. and Shelton, C. R · 2002
Cited alongside, same era.
Introduction to Rare Event Simulation
Bucklew, J. A · 2004
Cited alongside, same era.
On the markov chain central limit theorem
Jones, G. L. et al · 2004
Cited alongside, same era.
Conditional importance sampling estimators
Bucklew, J. A · 2005
Cited alongside, same era.
Approximate zero-variance simulation
L’Ecuyer, P. and Tuffin, B · 2008
Cited alongside, same era.
Importance sampling in rare event simulation
L’Ecuyer, P., Mandjes, M., and Tuffin, B · 2009
Cited alongside, same era.
Metelli, A. M., Papini, M., Faccio, F., and Restelli, M · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Sutton, R. and Barto, A · 2018
Later among the works it cites.
Off-policy deep reinforcement learning by bootstrapping the covariate shift
Gelada, C. and Bellemare, M. G · 2019
Closest in time.
Likelihood ratio gradient estimation for steady-state parameters
Glynn, P. W. and Olvera-Cravioto, M · 2019
Closest in time.
Empirical analysis of off-policy policy evaluation for reinforcement learning
Voloshin, C., Le, H. M., and Yue, Y · 2019
Closest in time.
Optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
Xie, T., Ma, Y., and Wang, Y · 2019
Closest in time.
Conditional importance sampling for off-policy learning
Rowland, M., Harutyunyan, A., van Hasselt, H., Borsa, D., Schaul, T., Munos, R., and Dabney, W · 2020
Closest in time.