Fetching the paper…
Reading the bibliography…
Reinforcement learning algorithms have had tremendous successes in online learning settings.
Benchmarking batch deep reinforcement learning algorithms
Scott Fujimoto, Edoardo Conti, Mohammad Ghavamzadeh, and Joelle Pineau · 1910
Earlier work this paper cites.
A Tutorial on Linear Function Approximators for Dynamic Programming and Reinforcement Learning
Alborz Geramifard, Thomas J. Walsh, Stefanie Tellex, Girish Chowdhary, Nicholas Roy, and Jonathan P. How · 1935
Earlier work this paper cites.
Methods of reducing sample size in monte carlo computations
Herman Kahn and Andy W Marshall · 1953
Earlier work this paper cites.
A markovian decision process
Richard Bellman · 1957
Earlier work this paper cites.
Structural equation models: An overview
Arthur S Goldberger · 1973
Earlier work this paper cites.
On the application of probability theory to agricultural experiments. essay on principles. section 9
Jerzy Splawa-Neyman, Dorota M Dabrowska, and TP Speed · 1990
Earlier work this paper cites.
Causal diagrams for empirical research
Judea Pearl · 1995
Earlier work this paper cites.
Temporal difference learning and td-gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
Integral probability metrics and their generating classes of functions
Alfred Müller · 1997
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup, Richard S Sutton, and Satinder Singh · 2000
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
Large-sample learning of bayesian networks is np-hard
David Maxwell Chickering, David Heckerman, and Christopher Meek · 2004
Earlier work this paper cites.
Stratification and weighting via the propensity score in estimation of causal treatment effects: a comparative study
Jared K Lunceford and Marie Davidian · 2004
Earlier work this paper cites.
Learning bayesian networks , volume 38
Richard E Neapolitan et al · 2004
Earlier work this paper cites.
Causal inference using potential outcomes: Design, modeling, decisions
Donald B Rubin · 2005
Earlier work this paper cites.
Demystifying double robustness: A comparison of alternative strategies for estimating a population mean from incomplete data
Joseph DY Kang, Joseph L Schafer, et al · 2007
Earlier work this paper cites.
Bias and variance approximation in value function estimates
Shie Mannor, Duncan Simester, Peng Sun, and John N Tsitsiklis · 2007
Earlier work this paper cites.
Estimation of causal effects using linear non-gaussian causal models with hidden variables
Patrik O Hoyer, Shohei Shimizu, Antti J Kerminen, and Markus Palviainen · 2008
Earlier work this paper cites.
The neyman-rubin model of causal inference and estimation via matching methods
Jasjeet S Sekhon · 2008
Earlier work this paper cites.
Domain adaptation: Learning bounds and algorithms
Yishay Mansour, Mehryar Mohri, and Afshin Rostamizadeh · 2009
Earlier work this paper cites.
Causal inference in statistics: An overview
Judea Pearl et al · 2009
Earlier work this paper cites.
Nonparametric return distribution approximation for reinforcement learning
Tetsuro Morimura, Masashi Sugiyama, Hisashi Kashima, Hirotaka Hachiya, and Toshiyuki Tanaka · 2010
Earlier work this paper cites.
Doubly robust policy evaluation and learning
Miroslav Dudík, John Langford, and Lihong Li · 2011
Earlier work this paper cites.
Doubly robust estimation of causal effects
Michele Jonsson Funk, Daniel Westreich, Chris Wiesen, Til Stürmer, M Alan Brookhart, and Marie Davidian · 2011
Cited alongside, same era.
The temporal logic of causal structures
Samantha Kleinberg and Bud Mishra · 2012
Cited alongside, same era.
Reinforcement learning in robotics: A survey
Jens Kober, J Andrew Bagnell, and Jan Peters · 2013
Cited alongside, same era.
Off-policy evaluation in Markov decision processes
Cosmin Paduraru · 2013
Cited alongside, same era.
Doubly robust policy evaluation and optimization
Miroslav Dudík, Dumitru Erhan, John Langford, Lihong Li, et al · 2014
Cited alongside, same era.
Introduction to structural equation models
Otis Dudley Duncan · 2014
Cited alongside, same era.
Estimating individual treatment effect: generalization bounds and algorithms
Uri Shalit, Fredrik D Johansson, and David Sontag · 2017
Later among the works it cites.
Transfer learning in multi-armed bandit: a causal approach
Junzhe Zhang and Elias Bareinboim · 2017
Later among the works it cites.
Woulda, coulda, shoulda: Counterfactually-guided policy search
Lars Buesing, Theophane Weber, Yori Zwols, Sebastien Racaniere, Arthur Guez, Jean-Baptiste Lespiau, and Nicolas Heess · 2018
Later among the works it cites.
Distributional reinforcement learning with quantile regression
Will Dabney, Mark Rowland, Marc G Bellemare, and Rémi Munos · 2018
Later among the works it cites.
More robust doubly robust off-policy evaluation
Mehrdad Farajtabar, Yinlam Chow, and Mohammad Ghavamzadeh · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Two optimal strategies for active learning of causal models from interventional data
Alain Hauser and Peter Bühlmann · 2014
Cited alongside, same era.
Machine learning methods for estimating heterogeneous causal effects
Susan Athey and Guido W Imbens · 2015
Cited alongside, same era.
Bandits with unobserved confounders: A causal approach
Elias Bareinboim, Andrew Forney, and Judea Pearl · 2015
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
High-confidence off-policy evaluation
Philip S Thomas, Georgios Theocharous, and Mohammad Ghavamzadeh · 2015
Cited alongside, same era.
The artificial intelligence clinician learns optimal treatment strategies for sepsis in intensive care
Matthieu Komorowski, Leo A Celi, Omar Badawi, Anthony C Gordon, and A Aldo Faisal · 2018
Later among the works it cites.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2018
Later among the works it cites.
Structural causal bandits: where to intervene?
Sanghack Lee and Elias Bareinboim · 2018
Later among the works it cites.
Representation balancing mdps for off-policy policy evaluation
Yao Liu, Omer Gottesman, Aniruddh Raghu, Matthieu Komorowski, Aldo A Faisal, Finale Doshi-Velez, and Emma Brunskill · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Striving for simplicity in off-policy deep reinforcement learning
Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi · 2019
Later among the works it cites.
Estimating treatment effects with causal forests: An application
Susan Athey and Stefan Wager · 2019
Later among the works it cites.
Human-level performance in 3d multiplayer games with population-based reinforcement learning
Max Jaderberg, Wojciech M Czarnecki, Iain Dunning, Luke Marris, Guy Lever, Antonio Garcia Castaneda, Charles Beattie, Neil C Rabinowitz, Ari S Morcos, Avraham Ruderman, et al · 2019
Later among the works it cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine · 2019
Later among the works it cites.
Safe policy improvement with baseline bootstrapping
Romain Laroche, Paul Trichelair, and Remi Tachet Des Combes · 2019
Later among the works it cites.
Structural causal bandits with non-manipulable variables
Sanghack Lee and Elias Bareinboim · 2019
Later among the works it cites.
A perspective on off-policy evaluation in reinforcement learning
Lihong Li · 2019
Later among the works it cites.
On convergence rate of adaptive multiscale value function approximation for reinforcement learning
Tao Li and Quanyan Zhu · 2019
Later among the works it cites.
Counterfactual off-policy evaluation with gumbel-max structural causal models
Michael Oberst and David Sontag · 2019
Later among the works it cites.
Adapting neural networks for the estimation of treatment effects
Claudia Shi, David Blei, and Victor Veitch · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Later among the works it cites.
Estimating treatment effects with observed confounders and mediators
Shantanu Gupta, Zachary C Lipton, and David Childers · 2020
Closest in time.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2062
Closest in time.