Fetching the paper…
Reading the bibliography…
We introduce an off-policy evaluation procedure for highlighting episodes where applying a reinforcement learned (RL) policy is likely to have produced a substantially different outcome than the observed policy.
Individual Choice Behavior: A Theoretical Analysis
Luce, R. D · 1959
Earlier work this paper cites.
Proportion of disease caused or prevented by a given exposure, trait or intervention
Miettinen, O. S · 1974
Earlier work this paper cites.
The relationship between Luce’s choice axiom, Thurstone’s theory of comparative judgment, and the double exponential distribution
Yellott, J. I · 1977
Earlier work this paper cites.
Counterfactual Probabilities: Computational Methods, Bounds and Applications
Balke, A. and Pearl, J · 1994
Earlier work this paper cites.
Probabilities of causation: three counterfactual interpretations and their identification
Pearl, J · 2000
Earlier work this paper cites.
Probabilities of causation : Bounds and identification
Tian, J. and Pearl, J · 2000
Earlier work this paper cites.
Discrete choice methods with simulation
Train, K · 2002
Earlier work this paper cites.
West’s Encyclopedia of American Law
Encyclopedia, W · 2008
Earlier work this paper cites.
Adaptive Treatment of Epilepsy via Batch-mode Reinforcement Learning
Guez, A., Vincent, R. D., Avoli, M., and Pineau, J · 2008
Earlier work this paper cites.
Dynamic discrete choice structural models: A survey
Aguirregabiria, V. and Mira, P · 2009
Earlier work this paper cites.
An introduction to medical malpractice in the United States
Bal, B. S · 2009
Earlier work this paper cites.
Causality: Models, Reasoning, and Inference
Pearl, J · 2009
Earlier work this paper cites.
Perturb-and-map random fields: Using discrete optimization to learn and sample from energy models
Yuille., G. P. and L, A · 2011
Earlier work this paper cites.
On the partition function and random maximum a-posteriori perturbations
Hazan, T. and Jaakkola, T · 2012
Cited alongside, same era.
Maddison, C. J., Tarlow, D., and Minka, T · 2014
Cited alongside, same era.
Causal Discovery with Continuous Additive Noise Models
Peters, J. and Schölkopf, B · 2014
Cited alongside, same era.
On the Causes of Effects: Response to Pearl
Dawid, P., Faigman, D. L., and Fienberg, S. E · 2015
Cited alongside, same era.
Causal Inference for Statistics, Social, and Biomedical Sciences
Imbens, G. W. and Rubin, D. B · 2015
Cited alongside, same era.
From statistical evidence to evidence of causality
Dawid, P., Musio, M., and Fienberg, S. E · 2016
Cited alongside, same era.
Gumbel Machinery, 2017
Maddison, C. J. and Tarlow, D · 2017
Later among the works it cites.
The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables
Maddison, C. J., Mnih, A., and Teh, Y. W · 2017
Later among the works it cites.
Combining Kernel and Model Based Learning for HIV Therapy Selection
Parbhoo, S., Bogojeska, J., Zazzi, M., Roth, V., and Doshi-Velez, F · 2017
Later among the works it cites.
Elements of Causal Inference: Foundations and Learning Algorithms
Peters, J., Janzing, D., and Schölkopf, B · 2017
Later among the works it cites.
Reliable Decision Support using Counterfactual Models
Schulam, P. and Saria, S · 2017
Later among the works it cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Perturbation, Optimization, and Statistics
Hazan, T., Papandreou, G., and Tarlow, D · 2016
Cited alongside, same era.
Learning Representations for Counterfactual Inference
Johansson, F. D., Shalit, U., and Sontag, D · 2016
Cited alongside, same era.
Distinguishing Cause from Effect Using Observational Data: Methods and Benchmarks
Mooij, J. M., Peters, J., Janzing, D., Zscheischler, J., and Schölkopf, B · 2016
Cited alongside, same era.
Data-Efficient Off-Policy Policy Evaluation for Reinforcement Learning
Thomas, P. S. and Brunskill, E · 2016
Cited alongside, same era.
Counterfactual-Based Prevented and Preventable Proportions
Yamada, K. and Kuroki, M · 2016
Cited alongside, same era.
Kocaoglu, M., Dimakis, A. G., Vishwanath, S., and Hassibi, B · 2017
Cited alongside, same era.
Later among the works it cites.
A nonparametric projection-based estimator for the probability of causation, with application to water sanitation in Kenya
Cuellar, M. and Kennedy, E. H · 2018
Later among the works it cites.
The Artificial Intelligence Clinician learns optimal treatment strategies for sepsis in intensive care
Komorowski, M., Celi, L. A., Badawi, O., Gordon, A. C., and Faisal, A. A · 2018
Later among the works it cites.
Structural Causal Bandits: Where to Intervene?
Lee, S. and Bareinboim, E · 2018
Later among the works it cites.
Representation Balancing MDPs for Off-policy Policy Evaluation
Liu, Y., Gottesman, O., Raghu, A., Komorowski, M., Faisal, A. A., Doshi-Velez, F., and Brunskill, E · 2018
Later among the works it cites.
Woulda, Coulda, Shoulda: Counterfactually-Guided Policy Search
Buesing, L., Weber, T., Zwols, Y., Heess, N., Racaniere, S., Guez, A., and Lespiau, J.-B · 2019
Closest in time.
Guidelines for reinforcement learning in healthcare
Gottesman, O., Johansson, F., Komorowski, M., Faisal, A., Sontag, D., Doshi-Velez, F., and Celi, L. A · 2019
Closest in time.