Fetching the paper…
Reading the bibliography…
We consider the problem of using expert data with unobserved confounders for imitation and reinforcement learning.
Alvinn: An autonomous land vehicle in a neural network
Dean A Pomerleau · 1989
Earlier work this paper cites.
Behavioural cloning: phenomena, results and problems
Ivan Bratko, Tanja Urbančič, and Claude Sammut · 1995
Earlier work this paper cites.
Causality: models, reasoning and inference , volume 29
Judea Pearl · 2000
Earlier work this paper cites.
Information theory and statistics: A tutorial
Imre Csiszár and Paul C Shields · 2004
Earlier work this paper cites.
On divergences and informations in statistics and information theory
Friedrich Liese and Igor Vajda · 2006
Earlier work this paper cites.
Calibrating sensitivity analyses to observed covariates in observational studies
Jesse Y Hsu and Dylan S Small · 2013
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Contextual markov decision processes
Assaf Hallak, Dotan Di Castro, and Shie Mannor · 2015
Earlier work this paper cites.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Earlier work this paper cites.
Causal bandits: learning good interventions via causal inference
Finnian Lattimore, Tor Lattimore, and Mark D Reid · 2016
Earlier work this paper cites.
Imitation learning: A survey of learning methods
Ahmed Hussein, Mohamed Medhat Gaber, Eyad Elyan, and Chrisina Jayne · 2017
Earlier work this paper cites.
Learning robust rewards with adversarial inverse reinforcement learning
Justin Fu, Katie Luo, and Sergey Levine · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Rllib: Abstractions for distributed reinforcement learning
Eric Liang, Richard Liaw, Robert Nishihara, Philipp Moritz, Roy Fox, Ken Goldberg, Joseph Gonzalez, Michael Jordan, and Ion Stoica · 2018
Earlier work this paper cites.
Learning robust rewards with adverserial inverse reinforcement learning
Justin Fu, Katie Luo, and Sergey Levine · 2018
Cited alongside, same era.
Imitation learning via kernel mean embedding
Kee-Eung Kim and Hyun Soo Park · 2018
Cited alongside, same era.
Nevergrad - A gradient-free optimization platform
J. Rapin and O. Teytaud · 2018
Cited alongside, same era.
Guidelines for reinforcement learning in healthcare
Omer Gottesman, Fredrik Johansson, Matthieu Komorowski, Aldo Faisal, David Sontag, Finale Doshi-Velez, and Leo Anthony Celi · 2019
Cited alongside, same era.
Recsim: A configurable simulation platform for recommender systems
Eugene Ie, Chih-wei Hsu, Martin Mladenov, Vihan Jain, Sanmit Narvekar, Jing Wang, Rui Wu, and Craig Boutilier · 2019
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
Accelerating online reinforcement learning with offline datasets
Ashvin Nair, Murtaza Dalal, Abhishek Gupta, and Sergey Levine · 2020
Later among the works it cites.
Causal imitation learning with unobserved confounders
Junzhe Zhang, Daniel Kumor, and Elias Bareinboim · 2020
Later among the works it cites.
Provably efficient causal reinforcement learning with confounded observational data
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ilya Kostrikov, Ofir Nachum, and Jonathan Tompson · 2019
Cited alongside, same era.
Disagreement-regularized imitation learning
Kianté Brantley, Wen Sun, and Mikael Henaff · 2019
Cited alongside, same era.
Algaedice: Policy gradient from arbitrary experience
Ofir Nachum, Bo Dai, Ilya Kostrikov, Yinlam Chow, Lihong Li, and Dale Schuurmans · 2019
Cited alongside, same era.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine · 2019
Cited alongside, same era.
Keep doing what worked: Behavior modelling priors for offline reinforcement learning
Noah Siegel, Jost Tobias Springenberg, Felix Berkenkamp, Abbas Abdolmaleki, Michael Neunert, Thomas Lampe, Roland Hafner, Nicolas Heess, and Martin Riedmiller · 2019
Cited alongside, same era.
Causal confusion in imitation learning
Pim de Haan, Dinesh Jayaraman, and Sergey Levine · 2019
Cited alongside, same era.
Near-optimal reinforcement learning in dynamic treatment regimes
Junzhe Zhang and Elias Bareinboim · 2019
Cited alongside, same era.
Lingxiao Wang, Zhuoran Yang, and Zhaoran Wang · 2020
Later among the works it cites.
Off-policy evaluation in partially observable environments
Guy Tennenholtz, Uri Shalit, and Shie Mannor · 2020
Later among the works it cites.
Confounding-robust policy evaluation in infinite-horizon reinforcement learning
Nathan Kallus and Angela Zhou · 2020
Later among the works it cites.
Imitation learning as f-divergence minimization
Liyiming Ke, Sanjiban Choudhury, Matt Barnes, Wen Sun, Gilwoo Lee, and Siddhartha Srinivasa · 2020
Later among the works it cites.
Off-policy policy evaluation for sequential decisions under unobserved confounding
Hongseok Namkoong, Ramtin Keramati, Steve Yadlowsky, and Emma Brunskill · 2020
Later among the works it cites.
A calculus for stochastic interventions: Causal effect identification and surrogate experiments
Juan Correa and Elias Bareinboim · 2020
Later among the works it cites.
Reward is enough for convex mdps
Tom Zahavy, Brendan O’Donoghue, Guillaume Desjardins, and Satinder Singh · 2021
Closest in time.
Offline reinforcement learning with fisher divergence critic regularization
Ilya Kostrikov, Rob Fergus, Jonathan Tompson, and Ofir Nachum · 2021
Closest in time.
A minimalist approach to offline reinforcement learning
Scott Fujimoto and Shixiang Shane Gu · 2021
Closest in time.
Minimax-optimal policy learning under unobserved confounding
Nathan Kallus and Angela Zhou · 2021
Closest in time.