Fetching the paper…
Reading the bibliography…
Empowered by expressive function approximators such as neural networks, deep reinforcement learning (DRL) achieves tremendous empirical successes.
Causal mediation analysis for stochastic interventions
Díaz, I · 1901
Earlier work this paper cites.
Reinforcement leaning in feature space: Matrix bandit, kernels, and regret bound
Yang, L. F · 1905
Earlier work this paper cites.
Provably efficient reinforcement learning with linear function approximation
Jin, C · 1907
Earlier work this paper cites.
Off-policy evaluation in partially observable environments
Tennenholtz, G · 1909
Earlier work this paper cites.
Regret analysis of causal bandit problems
Lu, Y · 1910
Earlier work this paper cites.
Provably efficient exploration in policy optimization
Cai, Q · 1912
Earlier work this paper cites.
Nonparametric bounds on treatment effects
Manski, C. F · 1990
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Bradtke, S. J · 1996
Earlier work this paper cites.
Li, C · 2003
Earlier work this paper cites.
Optimal dynamic treatment regimes
Murphy, S. A · 2003
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S · 2005
Earlier work this paper cites.
A distributional approach for causal inference using propensity scores
Tan, Z · 2006
Earlier work this paper cites.
Logarithmic online regret bounds for undiscounted reinforcement learning
Auer, P · 2007
Earlier work this paper cites.
Causality
Pearl, J · 2009
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Jaksch, T · 2010
Cited alongside, same era.
Introduction to the non-asymptotic analysis of random matrices
Vershynin, R · 2010
Cited alongside, same era.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y · 2011
Cited alongside, same era.
Population intervention causal effects based on stochastic interventions
Muñoz, I. D · 2012
Cited alongside, same era.
Counterfactuals and policy analysis in structural models
Balke, A · 2013
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Azar, M. G · 2017
Later among the works it cites.
Counterfactual data-fusion for online reinforcement learners
Forney, A · 2017
Later among the works it cites.
Elements of Causal Inference: Foundations and Learning Algorithms
Peters, J · 2017
Later among the works it cites.
Identifying best interventions through online importance sampling
Sen, R · 2017
Later among the works it cites.
Mastering the game of Go without human knowledge
Silver, D · 2017
Later among the works it cites.
Transfer learning in multi-armed bandit: A causal approach
Zhang, J · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforcement learning in robotics: Asurvey
Kober, J · 2013
Cited alongside, same era.
Dynamic treatment regimes
Chakraborty, B · 2014
Cited alongside, same era.
Generalization and exploration via randomized value functions
Osband, I · 2014
Cited alongside, same era.
Contextual Markov decision processes
Hallak, A · 2015
Cited alongside, same era.
Causal bandits: Learning good interventions via causal inference
Lattimore, F · 2016
Cited alongside, same era.
Deep reinforcement learning for dialogue generation
Li, J · 2016
Cited alongside, same era.
Buesing, L · 2018
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M · 2018
Later among the works it cites.
Is Q-learning provably efficient?
Jin, C · 2018
Later among the works it cites.
Deconfounding reinforcement learning in observational settings
Lu, C · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Sutton, R. S · 2018
Later among the works it cites.
Causal confusion in imitation learning
de Haan, P · 2019
Later among the works it cites.
Near-optimal reinforcement learning in dynamic treatment regimes
Zhang, J · 2019
Later among the works it cites.