Fetching the paper…
Reading the bibliography…
We introduce causal Markov Decision Processes (C-MDPs), a new formalism for sequential decision making which combines the standard MDP formulation with causal structures over state transition and reward functions.
Provably efficient reinforcement learning with linear function approximation
Jin, C., Yang, Z., Wang, Z., and Jordan, M. I. (2019) · 1907
Earlier work this paper cites.
Regret analysis of causal bandit problems
Lu, Y., Meisami, A., Tewari, A., and Yan, Z. (2019) · 1910
Earlier work this paper cites.
Causality: models, reasoning and inference
Pearl, J. (2000) · 2000
Earlier work this paper cites.
Off-policy policy evaluation for sequential decisions under unobserved confounding
Namkoong, H., Keramati, R., Yadlowsky, S., and Brunskill, E. (2020) · 2003
Earlier work this paper cites.
Probabilistic graphical models: principles and techniques
Koller, D. and Friedman, N. (2009) · 2009
Earlier work this paper cites.
Zhang, Z., Ji, X., and Du, S. S. (2020) · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P. (2010) · 2010
Earlier work this paper cites.
Regal: A regularization based algorithm for reinforcement learning in weakly communicating mdps
Bartlett, P. L. and Tewari, A. (2012) · 2012
Earlier work this paper cites.
Budgeted and non-budgeted causal bandits
Nair, V., Patil, V., and Sinha, G. (2020) · 2012
Cited alongside, same era.
(more) efficient reinforcement learning via posterior sampling
Osband, I., Russo, D., and Van Roy, B. (2013) · 2013
Cited alongside, same era.
Near-optimal reinforcement learning in factored mdps
Osband, I. and Van Roy, B. (2014) · 2014
Cited alongside, same era.
Causal bandits: Learning good interventions via causal inference
Lattimore, F., Lattimore, T., and Reid, M. D. (2016) · 2016
Cited alongside, same era.
Markov decision processes with unobserved confounders: A causal approach
Zhang, J. and Bareinboim, E. (2016) · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R. (2017) · 2017
Near-optimal reinforcement learning in dynamic treatment regimes
Zhang, J. and Bareinboim, E. (2019) · 2019
Later among the works it cites.
Regret minimization for reinforcement learning by evaluating the optimal bias function
Zhang, Z. and Ji, X. (2019) · 2019
Later among the works it cites.
Reinforcement learning for clinical decision support in critical care: comprehensive review
Liu, S., See, K. C., Ngiam, K. Y., Celi, L. A., Sun, X., and Feng, M. (2020) · 2020
Later among the works it cites.
Towards minimax optimal reinforcement learning in factored markov decision processes
Tian, Y., Qian, J., and Sra, S. (2020) · 2020
Later among the works it cites.
Is long horizon rl more difficult than short horizon rl?
Wang, R., Du, S. S., Yang, L., and Kakade, S. (2020) · 2020
Later among the works it cites.
Reinforcement learning in factored mdps: Oracle-efficient algorithms and tighter regret bounds for the non-episodic setting
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Identifying best interventions through online importance sampling
Sen, R., Shanmugam, K., Dimakis, A. G., and Shakkottai, S. (2017) · 2017
Cited alongside, same era.
Structural causal bandits: where to intervene?
Lee, S. and Bareinboim, E. (2018) · 2018
Cited alongside, same era.
Xu, Z. and Tewari, A. (2020) · 2020
Later among the works it cites.
Designing optimal dynamic treatment regimes: A causal reinforcement learning approach
Zhang, J. (2020) · 2020
Later among the works it cites.