Fetching the paper…
Reading the bibliography…
Off-policy evaluation (OPE) in reinforcement learning allows one to evaluate novel decision policies without needing to conduct exploration, which is often costly or otherwise infeasible.
A. Bibaut and M. van der Laan · 1907
Earlier work this paper cites.
A. F. Bibaut and M. J. van der Laan · 1907
Earlier work this paper cites.
Large sample properties of generalized method of moments estimators
L. P. Hansen · 1982
Earlier work this paper cites.
A new approach to causal inference in mortality studies with a sustained exposure period—application to control of the healthy worker survivor effect
J. Robins · 1986
Earlier work this paper cites.
Consistent estimation of the influence function of locally asymptotically linear estimators
C. A. J. Klaassen · 1987
Earlier work this paper cites.
On differentiable functionals
A. W. van der Vaart · 1991
Earlier work this paper cites.
Comment: Sequential moment restrictions in panel data
G. Chamberlain · 1992
Earlier work this paper cites.
Large sample estimation and hypothesis testing
W. K. Newey and D. L. Mcfadden · 1994
Earlier work this paper cites.
Rejoinder: The use of polynomial splines and their tensor products in multivariate function estimation
C. J. Stone · 1994
Earlier work this paper cites.
Finite-sample properties of some alternative gmm estimators
L. P. Hansen, J. Heaton, and A. Yaron · 1996
Earlier work this paper cites.
Efficient estimation of panel data models with sequential moment restrictions
J. Hahn · 1997
Earlier work this paper cites.
On methods of sieves and penalization
X. Shen · 1997
Earlier work this paper cites.
Efficient and Adaptive Estimation for Semiparametric Models
P. J. Bickel, C. A. J. Klaassen, Y. Ritov, and J. A. Wellner · 1998
Earlier work this paper cites.
On the role of the propensity score in efficient semiparametric estimation of average treatment effects
J. Hahn · 1998
Earlier work this paper cites.
Asymptotic statistics
A. W. van der Vaart · 1998
Earlier work this paper cites.
Adjusting for nonignorable dropout using semi-parametric models
D. Scharfstein, A. Rotnizky, and J. M. Robins · 1999
Earlier work this paper cites.
A matrix extension of the Cauchy-Schwarz inequality
G. Tripathi · 1999
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
D. Precup, R. Sutton, and S. Singh · 2000
Earlier work this paper cites.
Marginal structural models versus structural nested models as tools for causal inference
J. M. Robins · 2000
Earlier work this paper cites.
Marginal structural models and causal inference in epidemiology
J. M. Robins, M. A. Hernán, and B. Brumback · 2000
Earlier work this paper cites.
Marginal mean models for dynamic regimes
S. A. Murphy, M. J. van der Laan, and J. M. Robins · 2001
Earlier work this paper cites.
Semiparametric Statistics
A. W. van der Vaart · 2002
Earlier work this paper cites.
Efficient estimation of models with conditional moment restrictions containing unknown functions
C. Ai and X. Chen · 2003
Earlier work this paper cites.
Efficient estimation of average treatment effects using the estimated propensity score
K. Hirano, G. W. Imbens, and G. Ridder · 2003
Earlier work this paper cites.
Optimal dynamic treatment regimes
S. A. Murphy · 2003
Cited alongside, same era.
Unified Methods for Censored Longitudinal Data and Causality
M. J. van der Laan and J. M. Robins · 2003
Cited alongside, same era.
Least-squares policy iteration
M. Lagoudakis and R. Parr · 2004
Cited alongside, same era.
Local rademacher complexities
P. L. Bartlett, O. Bousquet, and S. Mendelson · 2005
Cited alongside, same era.
A distribution-free theory of nonparametric regression
L. Györfi, M. Kohler, A. Krzyzak, and H. Walk · 2006
Cited alongside, same era.
Semiparametric Theory and Missing Data
A. A. Tsiatis · 2006
Cited alongside, same era.
Chapter 76 large sample sieve estimation of semi-nonparametric models
The self-normalized estimator for counterfactual learning
A. Swaminathan and T. Joachims · 2015
Later among the works it cites.
The highly adaptive lasso estimator
D. Benkeser and M. van der Laan · 2016
Later among the works it cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, , and W. Zaremba · 2016
Later among the works it cites.
Doubly robust off-policy value evaluation for reinforcement learning
N. Jiang and L. Li · 2016
Later among the works it cites.
Safe and efficient off-policy reinforcement learning
R. Munos, T. Stepleton, A. Harutyunyan, and M. Bellemare · 2016
Later among the works it cites.
Data-efficient off-policy policy evaluation for reinforcement learning
P. Thomas and E. Brunskill · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. Chen · 2007
Cited alongside, same era.
Nonparametric econometrics : theory and practice
Q. Li and J. S. Racine · 2007
Cited alongside, same era.
Bias and variance approximation in value function estimates
S. Mannor, D. Simester, P. Sun, and J. N. Tsitsiklis · 2007
Cited alongside, same era.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
A. Antos, C. Szepesvári, and R. Munos · 2008
Cited alongside, same era.
Introduction to Empirical Processes and Semiparametric Inference
M. R. Kosorok · 2008
Cited alongside, same era.
Random features for large-scale kernel machines
A. Rahimi and B. Recht · 2008
Cited alongside, same era.
Later among the works it cites.
Adaptive concentration of regression trees, with application to random forests
S. Wager and G. Walther · 2016
Later among the works it cites.
Double/debiased machine learning for treatment and structural parameters
V. Chernozhukov, D. Chetverikov, M. Demirer, E. Duflo, C. Hansen, W. Newey, and J. Robins · 2018
Later among the works it cites.
Constructing dynamic treatment regimes over indefinite time horizons
A. Ertefaie and R. L. Strawderman · 2018
Later among the works it cites.
More robust doubly robust off-policy evaluation
M. Farajtabar, Y. Chow, and M. Ghavamzadeh · 2018
Later among the works it cites.
Deep neural networks learn non-smooth functions effectively
M. Imaizumi and K. Fukumizu · 2018
Later among the works it cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Q. Liu, L. Li, Z. Tang, and D. Zhou · 2018
Later among the works it cites.
Estimating dynamic treatment regimes in mobile health using v-learning
D. J. Luckett, E. B. Laber, A. R. Kahkoska, D. M. Maahs, E. Mayer-Davis, and M. R. Kosorok · 2018
Later among the works it cites.
Reinforcement learning : an introduction
R. S. Sutton · 2018
Later among the works it cites.
More efficient off-policy evaluation through regularized targeted learning
A. Bibaut, I. Malenica, N. Vlassis, and M. van der Laan · 2019
Closest in time.
Machine learning in the estimation of causal effects: targeted minimum loss-based estimation and double/debiased machine learning
I. n. Díaz · 2019
Closest in time.
Guidelines for reinforcement learning in healthcare
O. Gottesman, F. Johansson, M. Komorowski, A. Faisal, D. Sontag, F. Doshi-Velez, and L. A. Celi · 2019
Closest in time.
Causal Inference
M. Hernan and J. Robins · 2019
Closest in time.
Intrinsically efficient, stable, and bounded off-policy evaluation for reinforcement learning
N. Kallus and M. Uehara · 2019
Closest in time.
Non-parametric inference adaptive to intrinsic dimension
K. Khosravi, G. Lewis, and V. Syrgkanis · 2019
Closest in time.
Batch policy learning under constraints
H. Le, C. Voloshin, and Y. Yue · 2019
Closest in time.
Characterization of parameters with a mixed bias property
A. Rotnitzky, E. Smucler, and J. Robins · 2019
Closest in time.
Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
T. Xie, Y. Ma, and Y.-X. Wang · 2019
Closest in time.
Asymptotically efficient off-policy evaluation for tabular reinforcement learning
M. Yin and Y.-X. Wang · 2020
Closest in time.