Fetching the paper…
Reading the bibliography…
We propose a general formulation for addressing reinforcement learning (RL) problems in settings with observational data.
The design of experiments
Ronald Aylmer Fisher · 1935
Earlier work this paper cites.
Introduction to econometrics , volume 2
Gangadharrao Soundaryarao Maddala and Kajal Lahiri · 1992
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton, Andrew G Barto, et al · 1998
Earlier work this paper cites.
Measuring living standards with proxy variables
Mark R Montgomery, Michele Gragnolati, Kathleen A Burke, and Edmundo Paredes · 2000
Earlier work this paper cites.
Neural fitted q iteration–first experiences with a data efficient neural reinforcement learning method
Martin Riedmiller · 2005
Earlier work this paper cites.
Mostly harmless econometrics: An empiricist’s companion
Joshua D Angrist and Jörn-Steffen Pischke · 2008
Earlier work this paper cites.
Identifiability of parameters in latent structure models with many observed variables
Elizabeth S Allman, Catherine Matias, John A Rhodes, et al · 2009
Earlier work this paper cites.
Causality
Judea Pearl · 2009
Earlier work this paper cites.
On estimating firm-level production functions using proxy variables to control for unobservables
Jeffrey M Wooldridge · 2009
Earlier work this paper cites.
Learning individual and population level traits from clinical temporal data
Suchi Saria, Daphne Koller, and Anna Penn · 2010
Earlier work this paper cites.
Informing sequential clinical decision-making through reinforcement learning: an empirical study
Susan M Shortreed, Eric Laber, Daniel J Lizotte, T Scott Stroup, Joelle Pineau, and Susan A Murphy · 2011
Earlier work this paper cites.
On measurement bias in causal inference
Judea Pearl · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Developing predictive models using electronic medical records: challenges and pitfalls
Chris Paxton, Alexandru Niculescu-Mizil, and Suchi Saria · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Measurement bias and effect restoration in causal inference
Manabu Kuroki and Judea Pearl · 2014
Cited alongside, same era.
Stochastic backpropagation and approximate inference in deep generative models
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra · 2014
Cited alongside, same era.
Wojciech Zaremba and Ilya Sutskever · 2014
Cited alongside, same era.
Bandits with unobserved confounders: A causal approach
Elias Bareinboim, Andrew Forney, and Judea Pearl · 2015
Cited alongside, same era.
All your data are always missing: incorporating bias due to measurement error into the potential outcomes framework
Jessie K Edwards, Stephen R Cole, and Daniel Westreich · 2015
Cited alongside, same era.
Structured inference networks for nonlinear state space models
Rahul G Krishnan, Uri Shalit, and David Sontag · 2017
Later among the works it cites.
Causal effect inference with deep latent-variable models
Christos Louizos, Uri Shalit, Joris M Mooij, David Sontag, Richard Zemel, and Max Welling · 2017
Later among the works it cites.
Elements of causal inference: foundations and learning algorithms
Jonas Peters, Dominik Janzing, and Bernhard Schölkopf · 2017
Later among the works it cites.
A reinforcement learning approach to weaning of mechanical ventilation in intensive care units
Niranjani Prasad, Li-Fang Cheng, Corey Chivers, Michael Draugelis, and Barbara E Engelhardt · 2017
Later among the works it cites.
A causal multi-armed bandit approach for domestic robots’ failure avoidance
Nathan Ramoly, Amel Bouzeghoub, and Beatrice Finance · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rahul G Krishnan, Uri Shalit, and David Sontag · 2015
Cited alongside, same era.
The demise of early goal-directed therapy for severe sepsis and septic shock
PE Marik · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
The variational gaussian process
Dustin Tran, Rajesh Ranganath, and David M Blei · 2015
Cited alongside, same era.
Tensorflow: a system for large-scale machine learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al · 2016
Cited alongside, same era.
Onur Atan, William R Zame, Qiaojun Feng, and Mihaela van der Schaar · 2016
Cited alongside, same era.
Openai gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Reliable decision support using counterfactual models
Peter Schulam and Suchi Saria · 2017
Later among the works it cites.
Hossein Soleimani, Adarsh Subbaswamy, and Suchi Saria · 2017
Later among the works it cites.
Transfer learning in multi-armed bandit: a causal approach
Junzhe Zhang and Elias Bareinboim · 2017
Later among the works it cites.
Bayesian nonparametric causal inference: Information rates and learning algorithms
Ahmed M Alaa and Mihaela van der Schaar · 2018
Closest in time.
(non-)identification in latent confounder models
Alexander D’Amour · 2018
Closest in time.
Evaluating reinforcement learning algorithms in observational health settings
Omer Gottesman, Fredrik Johansson, Joshua Meier, Jack Dent, Donghun Lee, Srivatsan Srinivasan, Linying Zhang, Yi Ding, David Wihl, Xuefeng Peng, et al · 2018
Closest in time.
David Ha and Jürgen Schmidhuber · 2018
Closest in time.
Causal Inference
MA Hernán and JM Robins · 2018
Closest in time.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Sergey Levine · 2018
Closest in time.
Identifying causal effects with proxy variables of an unmeasured confounder
Wang Miao, Zhi Geng, and Eric J Tchetgen Tchetgen · 2018
Closest in time.
The Book of Why
Judea Pearl and Dana Mackenzie · 2018
Closest in time.
Woulda, coulda, shoulda: Counterfactually-guided policy search
Lars Buesing, Theophane Weber, Yori Zwols, Nicolas Heess, Sebastien Racaniere, Arthur Guez, and Jean-Baptiste Lespiau · 2019
Closest in time.