Fetching the paper…
Reading the bibliography…
In applications of offline reinforcement learning to observational data, such as in healthcare or education, a general concern is that observed actions might be affected by unobserved factors, inducing confounding and biasing estimates derived under the assumption of a perfect Markov decision process (MDP) model.
Large sample properties of generalized method of moments estimators
L. P. Hansen · 1982
Earlier work this paper cites.
On differentiable functionals
A. Van Der Vaart · 1991
Earlier work this paper cites.
Asymptotic statistics , volume 3
A. W. Van der Vaart · 2000
Earlier work this paper cites.
Linear inverse problems in structural econometrics estimation based on spectral decomposition and regularization
M. Carrasco, J.-P. Florens, and E. Renault · 2007
Earlier work this paper cites.
Efficient estimation of semiparametric conditional moment models with possibly nonsmooth residuals
X. Chen and D. Pouzo · 2009
Earlier work this paper cites.
Identification and estimation by penalization in nonparametric instrumental regression
J.-P. Florens, J. Johannes, and S. Van Bellegem · 2011
Earlier work this paper cites.
Cross-validated targeted minimum-loss-based estimation
W. Zheng and M. J. van der Laan · 2011
Earlier work this paper cites.
Estimation of nonparametric conditional moment models with possibly nonsmooth generalized residuals
X. Chen and D. Pouzo · 2012
Earlier work this paper cites.
Reinforcement learning of pomdps using spectral methods
K. Azizzadenesheli, A. Lazaric, and A. Anandkumar · 2016
Earlier work this paper cites.
Double machine learning for treatment and causal parameters
V. Chernozhukov, D. Chetverikov, M. Demirer, E. Duflo, C. Hansen, and W. K. Newey · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
D. Hendrycks and K. Gimpel · 2016
Earlier work this paper cites.
Learning in pomdps with monte carlo tree search
S. Katt, F. A. Oliehoek, and C. Amato · 2017
Earlier work this paper cites.
Counterfactual off-policy evaluation with gumbel-max structural causal models
M. Oberst and D. Sontag · 2019
Earlier work this paper cites.
Reinforcement learning for pomdp: Partitioned rollout and policy iteration with application to autonomous sequential repair problems
S. Bhattacharya, S. Badyal, T. Wheeler, S. Gil, and D. Bertsekas · 2020
Earlier work this paper cites.
Semiparametric proximal causal inference
Y. Cui, H. Pu, X. Shi, W. Miao, and E. J. Tchetgen Tchetgen · 2020
Cited alongside, same era.
Minimax estimation of conditional moment models
N. Dikkala, G. Lewis, L. Mackey, and V. Syrgkanis · 2020
Cited alongside, same era.
Double reinforcement learning for efficient off-policy evaluation in markov decision processes
N. Kallus and M. Uehara · 2020
Cited alongside, same era.
Confounding-robust policy evaluation in infinite-horizon reinforcement learning
N. Kallus and A. Zhou · 2020
Cited alongside, same era.
Provably efficient neural estimation of structural equation models: An adversarial approach
L. Liao, Y.-L. Chen, Z. Yang, B. Dai, M. Kolar, and Z. Wang · 2020
Cited alongside, same era.
Instrumental variable value iteration for causal offline reinforcement learning
L. Liao, Z. Fu, Z. Yang, M. Kolar, and Z. Wang · 2021
Closest in time.
A spectral approach to off-policy evaluation for pomdps
Y. Nair and N. Jiang · 2021
Closest in time.
Structured world belief for reinforcement learning in pomdp
G. Singh, S. Peri, J. Kim, H. Kim, and S. Ahn · 2021
Closest in time.
Provably efficient causal reinforcement learning with confounded observational data
L. Wang, Z. Yang, and Z. Wang · 2021
Closest in time.
Deep proxy causal learning and its application to confounded bandit policy evaluation
L. Xu, H. Kanagawa, and A. Gretton · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Namkoong, R. Keramati, S. Yadlowsky, and E. Brunskill · 2020
Cited alongside, same era.
Imitation learning for agile autonomous driving
Y. Pan, C.-A. Cheng, K. Saigol, K. Lee, X. Yan, E. A. Theodorou, and B. Boots · 2020
Cited alongside, same era.
Multiply robust causal inference with double-negative control adjustment for categorical unmeasured confounding
X. Shi, W. Miao, J. C. Nelson, and E. J. Tchetgen Tchetgen · 2020
Cited alongside, same era.
An introduction to proximal causal learning
E. J. Tchetgen Tchetgen, A. Ying, Y. Cui, X. Shi, and W. Miao · 2020
Cited alongside, same era.
Off-policy evaluation in partially observable environments
G. Tennenholtz, S. Mannor, and U. Shalit · 2020
Cited alongside, same era.
Off-policy evaluation in infinite-horizon reinforcement learning with latent confounders
A. Bennett, N. Kallus, L. Li, and A. Mousavi · 2021
Cited alongside, same era.
Universal off-policy evaluation
Y. Chandak, S. Niekum, B. da Silva, E. Learned-Miller, E. Brunskill, and P. S. Thomas · 2021
Cited alongside, same era.
C.-H. H. Yang, I. Hung, T. Danny, Y. Ouyang, and P.-Y. Chen · 2021
Closest in time.
Proximal causal inference for complex longitudinal studies
A. Ying, W. Miao, X. Shi, and E. J. Tchetgen Tchetgen · 2021
Closest in time.
Minimax kernel machine learning for a class of doubly robust functionals with application to proximal causal inference
A. Ghassami, A. Ying, I. Shpitser, and E. T. Tchetgen · 2022
Closest in time.
Efficiently breaking the curse of horizon in off-policy evaluation with double reinforcement learning
N. Kallus and M. Uehara · 2022
Closest in time.
Causal inference under unmeasured confounding with negative controls: A minimax learning approach
N. Kallus, X. Mao, and M. Uehara · 2022
Closest in time.
Counterfactually guided policy transfer in clinical settings
T. W. Killian, M. Ghassemi, and S. Joshi · 2022
Closest in time.
Spectral representation learning for conditional moment models
Z. Wang, Y. Luo, Y. Li, J. Zhu, and B. Schölkopf · 2022
Closest in time.
The variational method of moments
A. Bennett and N. Kallus · 2023
Closest in time.
Off-policy evaluation in partially observed markov decision processes under sequential ignorability
Y. Hu and S. Wager · 2023
Closest in time.