Fetching the paper…
Reading the bibliography…
In offline reinforcement learning (RL) an optimal policy is learned solely from a priori collected observational data.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Linear Integral Equations
R. Kress · 1989
Earlier work this paper cites.
Nonparametric bounds on treatment effects
C. F. Manski · 1990
Earlier work this paper cites.
Counterfactual probabilities: Computational methods, bounds and applications
A. Balke and J. Pearl · 1994
Earlier work this paper cites.
Bounds on treatment effects from studies with imperfect compliance
A. Balke and J. Pearl · 1997
Earlier work this paper cites.
Observational Studies
P. Rosenbaum · 2002
Earlier work this paper cites.
Optimal dynamic treatment regimes
S. A. Murphy · 2003
Earlier work this paper cites.
Instrumental variable estimation of nonparametric models
W. K. Newey and J. L. Powell · 2003
Earlier work this paper cites.
An IV model of quantile treatment effects
V. Chernozhukov and C. Hansen · 2005
Earlier work this paper cites.
Semi-nonparametric IV estimation of shape-invariant Engel curves
R. Blundell, X. Chen, and D. Kristensen · 2007
Earlier work this paper cites.
Preference-based instrumental variable methods for the estimation of treatment effects: Assessing validity and interpreting results
M. A. Brookhart and S. Schneeweiss · 2007
Earlier work this paper cites.
Linear inverse problems in structural econometrics estimation based on spectral decomposition and regularization
M. Carrasco, J.-P. Florens, and E. Renault · 2007
Earlier work this paper cites.
Instrumental variable estimation of nonseparable models
V. Chernozhukov, G. W. Imbens, and W. K. Newey · 2007
Earlier work this paper cites.
Nonparametric instrumental variables estimation of a quantile regression model
J. L. Horowitz and S. Lee · 2007
Earlier work this paper cites.
Mostly harmless econometrics: An empiricist’s companion
J. D. Angrist and J.-S. Pischke · 2008
Earlier work this paper cites.
Adaptive treatment of epilepsy via batch-mode reinforcement learning
A. Guez, R. D. Vincent, M. Avoli, and J. Pineau · 2008
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Earlier work this paper cites.
Causality
J. Pearl · 2009
Earlier work this paper cites.
Confounding control in healthcare database research: Challenges and potential approaches
M. A. Brookhart, T. Stürmer, R. J. Glynn, J. Rassen, and S. Schneeweiss · 2010
Earlier work this paper cites.
A first-order primal-dual algorithm for convex problems with applications to imaging
A. Chambolle and T. Pock · 2011
Earlier work this paper cites.
Nonparametric instrumental regression
S. Darolles, Y. Fan, J.-P. Florens, and E. Renault · 2011
Earlier work this paper cites.
On the completeness condition in nonparametric instrumental problems
X. D’Haultfoeuille · 2011
Earlier work this paper cites.
Causal inference by surrogate experiments: z -identifiability
E. Bareinboim and J. Pearl · 2012
Earlier work this paper cites.
The differential impact of delivery hospital on the outcomes of premature infants
S. A. Lorch, M. Baiocchi, C. E. Ahlberg, and D. S. Small · 2012
Earlier work this paper cites.
Statistical Methods for Dynamic Treatment Regimes
B. Chakraborty and E. E. Moodie · 2013
Earlier work this paper cites.
Instrumental variable methods for causal inference
M. Baiocchi, J. Cheng, and D. S. Small · 2014
Earlier work this paper cites.
Dynamic treatment regimes
B. Chakraborty and S. A. Murphy · 2014
Earlier work this paper cites.
The movielens datasets: History and context
F. M. Harper and J. A. Konstan · 2015
Earlier work this paper cites.
MIMIC-III, a freely accessible critical care database
A. E. Johnson, T. J. Pollard, L. Shen, H. L. Li-Wei, M. Feng, M. Ghassemi, B. Moody, P. Szolovits, L. A. Celi, and R. G. Mark · 2016
Cited alongside, same era.
Markov decision processes with unobserved confounders: A causal approach
J. Zhang and E. Bareinboim · 2016
Cited alongside, same era.
Examples of L2-complete and boundedly-complete distributions
D. W. Andrews · 2017
Cited alongside, same era.
Learning from conditional distributions via dual embeddings
B. Dai, N. He, Y. Pan, B. Boots, and L. Song · 2017
Cited alongside, same era.
Stochastic variance reduction methods for policy evaluation
S. S. Du, J. Chen, L. Li, L. Xiao, and D. Zhou · 2017
Cited alongside, same era.
Nonparametric identification using instrumental variables: Sufficient conditions for completeness
Y. Hu and J.-L. Shiu · 2017
Kernel instrumental variable regression
R. Singh, M. Sahani, and A. Gretton · 2019
Later among the works it cites.
Near-optimal reinforcement learning in dynamic treatment regimes
J. Zhang and E. Bareinboim · 2019
Later among the works it cites.
Provably efficient exploration in policy optimization
Q. Cai, Z. Yang, C. Jin, and Z. Wang · 2020
Later among the works it cites.
Minimax estimation of conditional moment models
N. Dikkala, G. Lewis, L. Mackey, and V. Syrgkanis · 2020
Later among the works it cites.
Information theoretic regret bounds for online nonlinear control
S. Kakade, A. Krishnamurthy, K. Lowrey, M. Ohnishi, and W. Sun · 2020
Later among the works it cites.
Confounding-robust policy evaluation in infinite-horizon reinforcement learning
N. Kallus and A. Zhou · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Combining kernel and model based learning for HIV therapy selection
S. Parbhoo, J. Bogojeska, M. Zazzi, V. Roth, and F. Doshi-Velez · 2017
Cited alongside, same era.
Elements of Causal Inference: Foundations and Learning Algorithms
J. Peters, D. Janzing, and B. Schölkopf · 2017
Cited alongside, same era.
A reinforcement learning approach to weaning of mechanical ventilation in intensive care units
N. Prasad, L.-F. Cheng, C. Chivers, M. Draugelis, and B. E. Engelhardt · 2017
Cited alongside, same era.
Continuous state-space models for optimal sepsis treatment: a deep reinforcement learning approach
A. Raghu, M. Komorowski, L. A. Celi, P. Szolovits, and M. Ghassemi · 2017
Cited alongside, same era.
Exploiting strong convexity from data with primal-dual first-order algorithms
J. Wang and L. Xiao · 2017
Cited alongside, same era.
Woulda, coulda, shoulda: counterfactually-guided policy search
L. Buesing, T. Weber, Y. Zwols, S. Racaniere, A. Guez, J.-B. Lespiau, and N. Heess · 2018
Cited alongside, same era.
Provably efficient neural estimation of structural equation model: An adversarial approach
L. Liao, Y.-L. Chen, Y. Zhuoran, B. Dai, M. Kolar, and Z. Wang · 2020
Later among the works it cites.
Dual IV: A single stage instrumental variable regression
K. Muandet, A. Mehrjou, S. K. Lee, and A. Raj · 2020
Later among the works it cites.
Off-policy policy evaluation for sequential decisions under unobserved confounding
H. Namkoong, R. Keramati, S. Yadlowsky, and E. Brunskill · 2020
Later among the works it cites.
Off-policy evaluation in partially observable environments
G. Tennenholtz, U. Shalit, and S. Mannor · 2020
Later among the works it cites.
Causal inference for recommender systems
Y. Wang, D. Liang, L. Charlin, and D. M. Blei · 2020
Later among the works it cites.
Designing optimal dynamic treatment regimes: A causal reinforcement learning approach
J. Zhang and E. Bareinboim · 2020
Later among the works it cites.
Off-policy evaluation in infinite-horizon reinforcement learning with latent confounders
A. Bennett, N. Kallus, L. Li, and A. Mousavi · 2021
Closest in time.
A semiparametric instrumental variable approach to optimal treatment regimes under endogeneity
Y. Cui and E. Tchetgen Tchetgen · 2021
Closest in time.
Minimax-optimal policy learning under unobserved confounding
N. Kallus and A. Zhou · 2021
Closest in time.
Estimating optimal treatment rules with an instrumental variable: A partial identification learning approach
H. Pu and B. Zhang · 2021
Closest in time.
Provably efficient causal reinforcement learning with confounded observational data
L. Wang, Z. Yang, and Z. Wang · 2021
Closest in time.
Bounding causal effects on continuous outcomes
J. Zhang and E. Bareinboim · 2021
Closest in time.
Adaptively exploiting d-separators with causal bandits
B. Bilodeau, L. Wang, and D. Roy · 2022
Closest in time.
Offline reinforcement learning with instrumental variables in confounded Markov decision processes
Z. Fu, Z. Qi, Z. Wang, Z. Yang, Y. Xu, and M. R. Kosorok · 2022
Closest in time.
Active learning for nonlinear system identification with guarantees
H. Mania, M. I. Jordan, and B. Recht · 2022
Closest in time.
A free lunch from the noise: Provable and practical exploration for representation learning
T. Ren, T. Zhang, C. Szepesvári, and B. Dai · 2022
Closest in time.
M. Yu, Z. Yang, and J. Fan · 2022
Closest in time.
Proximal reinforcement learning: Efficient off-policy evaluation in partially observed Markov decision processes
A. Bennett and N. Kallus · 2023
Closest in time.
Estimating and improving dynamic treatment regimes with a time-varying instrumental variable
S. Chen and B. Zhang · 2023
Closest in time.
Causal inference and data fusion in econometrics
P. Hünermund and E. Bareinboim · 2023
Closest in time.
Instrumental variable estimation of marginal structural mean models for time-varying treatment
H. Michael, Y. Cui, S. A. Lorch, and E. J. Tchetgen Tchetgen · 2023
Closest in time.
An instrumental variable approach to confounded off-policy evaluation
Y. Xu, J. Zhu, C. Shi, S. Luo, and R. Song · 2023
Closest in time.