Fetching the paper…
Reading the bibliography…
When observed decisions depend only on observed features, off-policy policy evaluation (OPE) methods for sequential decision making problems can estimate the performance of evaluation policies before deploying them.
Smoking and lung cancer: Recent evidence and a discussion of some questions
J. Cornfield, W. Haenszel, E. C. Hammond, A. M. Lilienfeld, M. B. Shimkin, and E. L. Wynder · 1959
Earlier work this paper cites.
A five to fifteen year follow-up study of infantile psychosis: Ii. social and behavioural outcome
M. Rutter, D. Greenfeld, and L. Lockyer · 1967
Earlier work this paper cites.
Optimization by Vector Space Methods
D. Luenberger · 1969
Earlier work this paper cites.
Assessing sensitivity to an unobserved binary covariate in an observational study with binary outcome
P. R. Rosenbaum and D. B. Rubin · 1983
Earlier work this paper cites.
A new approach to causal inference in mortality studies with a sustained exposure period—application to control of the healthy worker survivor effect
J. Robins · 1986
Earlier work this paper cites.
Strong laws of large numbers for arrays of rowwise independent random variables
T.-C. Hu, F. Moricz, and R. Taylor · 1989
Earlier work this paper cites.
Nonparametric bounds on treatment effects
C. F. Manski · 1990
Earlier work this paper cites.
Epi-consistency of convex stochastic programs
A. J. King and R. J. Wets · 1991
Earlier work this paper cites.
Medical heuristics: the silent adjudicators of clinical practice
C. J. McDonald · 1996
Earlier work this paper cites.
Causal inference from complex longitudinal data
J. M. Robins · 1997
Earlier work this paper cites.
Variational Analysis
R. T. Rockafellar and R. J. B. Wets · 1998
Earlier work this paper cites.
Sensitivity analysis for selection bias and unmeasured confounding in missing data and causal inference models
J. M. Robins, A. Rotnitzky, and D. O. Scharfstein · 2000
Earlier work this paper cites.
Marginal mean models for dynamic regimes
S. A. Murphy, M. J. van der Laan, and J. M. Robins · 2001
Earlier work this paper cites.
Observational studies
P. R. Rosenbaum · 2002
Earlier work this paper cites.
Psychological models of professional decision making
M. K. Dhami · 2003
Earlier work this paper cites.
Sensitivity to exogeneity assumptions in program evaluation
G. W. Imbens · 2003
Earlier work this paper cites.
Optimal dynamic treatment regimes
S. A. Murphy · 2003
Earlier work this paper cites.
Sensitivity analyses for unmeasured confounding assuming a marginal structural model for repeated measures
B. A. Brumback, M. A. Hernán, S. J. P. A. Haneuse, and J. M. Robins · 2004
Earlier work this paper cites.
Optimal structural nested models for optimal sequential decisions
J. M. Robins · 2004
Earlier work this paper cites.
The effectiveness of simple decision heuristics: Forecasting commercial success for early-stage ventures
T. Åstebro and S. Elhedhli · 2006
Cited alongside, same era.
Instant customer base analysis: Managerial heuristics often “get it right”
M. Wübben and F. v. Wangenheim · 2008
Cited alongside, same era.
Patterns of growth in adaptive social abilities among children with autism spectrum disorders
D. K. Anderson, R. S. Oti, C. Lord, and K. Welch · 2009
Cited alongside, same era.
Causality
J. Pearl · 2009
Cited alongside, same era.
Design of Observational Studies , volume 10
P. R. Rosenbaum · 2010
Cited alongside, same era.
Extraneous factors in judicial decisions
S. Danziger, J. Levav, and L. Avnaim-Pesso · 2011
Cited alongside, same era.
Surviving sepsis campaign: International guidelines for management of sepsis and septic shock: 2016
A. Rhodes, L. E. Evans, W. Alhazzani, et al · 2017
Later among the works it cites.
Time to treatment and mortality during mandated emergency care for sepsis
C. W. Seymour, F. Gesten, H. C. Prescott, M. E. Friedrich, T. J. Iwashyna, G. S. Phillips, S. Lemeshow, T. Osborn, K. M. Terry, and M. M. Levy · 2017
Later among the works it cites.
Learning to treat sepsis with multi-output gaussian process deep recurrent q-networks, 2018
J. Futoma, A. Lin, M. Sendak, A. Bedoya, M. Clement, C. O’Brien, and K. Heller · 2018
Later among the works it cites.
Algorithmic decision making in the presence of unmeasured confounding
J. Jung, R. Shroff, A. Feller, and S. Goel · 2018
Later among the works it cites.
Confounding-robust policy improvement
N. Kallus and A. Zhou · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Communication interventions for minimally verbal children with autism: A sequential multiple assignment randomized trial
C. Kasari, A. Kaiser, K. Goods, J. Nietfeld, P. Mathy, R. Landa, S. Murphy, and D. Almirall · 2014
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
N. Jiang and L. Li · 2015
Cited alongside, same era.
The impact of timing of antibiotics on outcomes in severe sepsis and septic shock: a systematic review and meta-analysis
S. A. Sterling, W. R. Miller, J. Pryor, M. A. Puskarich, and A. E. Jones · 2015
Cited alongside, same era.
High-confidence off-policy evaluation
P. S. Thomas, G. Theocharous, and M. Ghavamzadeh · 2015
Cited alongside, same era.
Making contextual decisions with low technical debt
A. Agarwal, S. Bird, M. Cozowicz, L. Hoang, J. Langford, S. Lee, J. Li, D. Melamed, G. Oshri, O. Ribas, et al · 2016
Cited alongside, same era.
Mimic-iii, a freely accessible critical care database
A. E. Johnson, T. J. Pollard, L. Shen, H. L. Li-wei, M. Feng, M. Ghassemi, B. Moody, P. Szolovits, L. A. Celi, and R. G. Mark · 2016
Cited alongside, same era.
N. Kallus, X. Mao, and A. Zhou · 2018
Later among the works it cites.
Bounds on the conditional and average treatment effect in the presence of unobserved confounders
S. Yadlowsky, H. Namkoong, S. Basu, J. Duchi, and L. Tian · 2018
Later among the works it cites.
Off-policy deep reinforcement learning without exploration
S. Fujimoto, D. Meger, and D. Precup · 2019
Later among the works it cites.
Importance sampling policy evaluation with an estimated behavior policy
J. Hanna, S. Niekum, and P. Stone · 2019
Later among the works it cites.
Double reinforcement learning for efficient off-policy evaluation in markov decision processes
N. Kallus and M. Uehara · 2019
Later among the works it cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
A. Kumar, J. Fu, M. Soh, G. Tucker, and S. Levine · 2019
Later among the works it cites.
Batch policy learning under constraints
H. M. Le, C. Voloshin, and Y. Yue · 2019
Later among the works it cites.
Off-policy policy gradient with state distribution correction
Y. Liu, A. Swaminathan, A. Agarwal, and E. Brunskill · 2019
Later among the works it cites.
Learning when-to-treat policies
X. Nie, E. Brunskill, and S. Wager · 2019
Later among the works it cites.
Counterfactual off-policy evaluation with Gumbel-max structural causal models
M. Oberst and D. Sontag · 2019
Later among the works it cites.
Preventing undesirable behavior of intelligent machines
P. S. Thomas, B. C. da Silva, A. G. Barto, S. Giguere, Y. Brun, and E. Brunskill · 2019
Later among the works it cites.
Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
T. Xie, Y. Ma, and Y.-X. Wang · 2019
Later among the works it cites.
Near-optimal reinforcement learning in dynamic treatment regimes
J. Zhang and E. Bareinboim · 2019
Later among the works it cites.
Causal Inference: What If
M. Hernán and J. Robins · 2020
Closest in time.