Fetching the paper…
Reading the bibliography…
While Reinforcement Learning (RL) achieves tremendous success in sequential decision-making problems of many domains, it still faces key challenges of data inefficiency and the lack of interpretability.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in International Conference on Machine Learning . PMLR, 2016, pp. 1928–1937
1937
Earlier work this paper cites.
R. Bellman, “Dynamic programming,” Science , vol. 153, no. 3731, pp. 34–37, 1966
1966
Earlier work this paper cites.
C. W. Granger, “Investigating causal relations by econometric models and cross-spectral methods,” Econometrica: journal of the Econometric Society , pp. 424–438, 1969
1969
Earlier work this paper cites.
D. B. Rubin, “Estimating causal effects of treatments in randomized and nonrandomized studies.” Journal of educational Psychology , vol. 66, no. 5, p. 688, 1974
1974
Earlier work this paper cites.
G. Schwarz, “Estimating the dimension of a model,” The annals of statistics , pp. 461–464, 1978
1978
Earlier work this paper cites.
P. R. Rosenbaum and D. B. Rubin, “The central role of the propensity score in observational studies for causal effects,” Biometrika , vol. 70, no. 1, pp. 41–55, 1983
1983
Earlier work this paper cites.
R. S. Sutton, “Integrated architectures for learning, planning, and reacting based on approximating dynamic programming,” in Machine learning proceedings 1990 . Elsevier, 1990, pp. 216–224
1990
Earlier work this paper cites.
P. Spirtes and C. Glymour, “An algorithm for fast recovery of sparse causal graphs,” Social science computer review , vol. 9, no. 1, pp. 62–72, 1991
1991
Earlier work this paper cites.
D. Heckerman, D. Geiger, and D. M. Chickering, “Learning bayesian networks: The combination of knowledge and statistical data,” Machine learning , vol. 20, no. 3, pp. 197–243, 1995
1995
Earlier work this paper cites.
R. S. Sutton, A. G. Barto et al. , Introduction to reinforcement learning . MIT press Cambridge, 1998
1998
Earlier work this paper cites.
V. Konda and J. Tsitsiklis, “Actor-critic algorithms,” Advances in Neural Information Processing Systems , vol. 12, 1999
1999
Earlier work this paper cites.
J. Pearl, Causality: Models, Reasoning, and Inference . New York, NY, USA: Cambridge University Press, 2000
2000
Earlier work this paper cites.
P. Spirtes, C. N. Glymour, and R. Scheines, Causation, Prediction, and Search . MIT press, 2000
2000
Earlier work this paper cites.
J. M. Robins, A. Rotnitzky, and D. O. Scharfstein, “Sensitivity analysis for selection bias and unmeasured confounding in missing data and causal inference models,” in Statistical models in epidemiology, the environment, and clinical trials . Springer, 2000, pp. 1–94
2000
Earlier work this paper cites.
D. M. Chickering, “Optimal structure identification with greedy search,” Journal of Machine Learning Research , vol. 3, no. Nov, pp. 507–554, 2002
2002
Earlier work this paper cites.
S. A. Murphy, “Optimal dynamic treatment regimes,” Journal of the Royal Statistical Society: Series B (Statistical Methodology) , vol. 65, no. 2, pp. 331–355, 2003
2003
Earlier work this paper cites.
Z. Wang, T. Schaul, M. Hessel, H. Hasselt, M. Lanctot, and N. Freitas, “Dueling network architectures for deep reinforcement learning,” in International Conference on Machine Learning . PMLR, 2016, pp. 1995–2003
2003
Earlier work this paper cites.
S. Shimizu, P. O. Hoyer, A. Hyvärinen, and A. Kerminen, “A linear non-Gaussian acyclic model for causal discovery,” Journal of Machine Learning Research , vol. 7, no. Oct, pp. 2003–2030, 2006
2006
Earlier work this paper cites.
Z. Tan, “A distributional approach for causal inference using propensity scores,” Journal of the American Statistical Association , vol. 101, no. 476, pp. 1619–1637, 2006
2006
Earlier work this paper cites.
J. Langford and T. Zhang, “The epoch-greedy algorithm for contextual multi-armed bandits,” Advances in Neural Information Processing Systems , vol. 20, no. 1, pp. 96–1, 2007
2007
Earlier work this paper cites.
A. Bennett, N. Kallus, L. Li, and A. Mousavi, “Off-policy evaluation in infinite-horizon reinforcement learning with latent confounders,” in International Conference on Artificial Intelligence and Statistics . PMLR, 2021, pp. 1999–2007
2007
Earlier work this paper cites.
J. S. Sekhon, “The neyman-rubin model of causal inference and estimation via matching methods,” The Oxford handbook of political methodology , vol. 2, pp. 1–32, 2008
2008
Earlier work this paper cites.
P. Hoyer, D. Janzing, J. M. Mooij, J. Peters, and B. Schölkopf, “Nonlinear causal discovery with additive noise models,” Advances in Neural Information Processing Systems , vol. 21, 2008
2008
Earlier work this paper cites.
A. Hyvärinen, S. Shimizu, and P. O. Hoyer, “Causal modelling combining instantaneous and lagged effects: an identifiable model based on non-gaussianity,” in Proceedings of the 25th International Conference on Machine Learning , 2008, pp. 424–431
2008
Earlier work this paper cites.
M. Caliendo and S. Kopeinig, “Some practical guidance for the implementation of propensity score matching,” Journal of economic surveys , vol. 22, no. 1, pp. 31–72, 2008
2008
Earlier work this paper cites.
G. Chaslot, S. Bakkes, I. Szita, and P. Spronck, “Monte-carlo tree search: A new framework for game ai,” in Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment , vol. 4, no. 1, 2008, pp. 216–217
2008
Earlier work this paper cites.
K. Zhang and A. Hyvärinen, “On the identifiability of the post-nonlinear causal model,” in Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence (UAI) . AUAI Press, 2009, pp. 647–655
2009
Earlier work this paper cites.
D. Entner and P. O. Hoyer, “On causal discovery from time series data using fci,” Probabilistic graphical models , pp. 121–128, 2010
2010
Earlier work this paper cites.
B. D. Ziebart, J. A. Bagnell, and A. K. Dey, “Modeling interaction via the principle of maximum causal entropy,” in International Conference on Machine Learning , 2010
2010
Earlier work this paper cites.
K. Zhang, J. Peters, D. Janzing, and B. Schölkopf, “Kernel-based conditional independence test and application in causal discovery,” in Proceedings of the Twenty-Seventh Conference on Uncertainty in Artificial Intelligence , 2011, pp. 804–813
2011
Earlier work this paper cites.
J. L. Hill, “Bayesian nonparametric modeling for causal inference,” Journal of Computational and Graphical Statistics , vol. 20, no. 1, pp. 217–240, 2011
2011
Earlier work this paper cites.
M. Deisenroth and C. E. Rasmussen, “Pilco: A model-based and data-efficient approach to policy search,” in Proceedings of the 28th International Conference on Machine Learning . Citeseer, 2011, pp. 465–472
2011
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” in 2012 IEEE/RSJ international conference on intelligent robots and systems . IEEE, 2012, pp. 5026–5033
2012
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller, “Playing atari with deep reinforcement learning,” Advances in Neural Information Processing Systems (NIPS) , 2013
2013
Earlier work this paper cites.
J. Kober, J. A. Bagnell, and J. Peters, “Reinforcement learning in robotics: A survey,” The International Journal of Robotics Research , vol. 32, no. 11, pp. 1238–1274, 2013
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
S. Levine and V. Koltun, “Guided policy search,” in International Conference on Machine Learning . PMLR, 2013, pp. 1–9
2013
Earlier work this paper cites.
——, “The principle of maximum causal entropy for estimating interacting processes,” IEEE Transactions on Information Theory , vol. 59, no. 4, pp. 1966–1980, 2013
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
M. P. Deisenroth, G. Neumann, J. Peters et al. , “A survey on policy search for robotics,” Foundations and Trends® in Robotics , vol. 2, no. 1–2, pp. 1–142, 2013
2013
Earlier work this paper cites.
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller, “Deterministic policy gradient algorithms,” in International Conference on Machine Learning . PMLR, 2014, pp. 387–395
2014
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al. , “Human-level control through deep reinforcement learning,” nature , vol. 518, no. 7540, pp. 529–533, 2015
2015
Earlier work this paper cites.
E. Bareinboim, A. Forney, and J. Pearl, “Bandits with unobserved confounders: A causal approach,” Advances in Neural Information Processing Systems , vol. 28, 2015
2015
Earlier work this paper cites.
G. W. Imbens and D. B. Rubin, Causal inference in statistics, social, and biomedical sciences . Cambridge University Press, 2015
2015
Earlier work this paper cites.
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, “Trust region policy optimization,” in International Conference on Machine Learning . PMLR, 2015, pp. 1889–1897
2015
Earlier work this paper cites.
C. Finn, S. Levine, and P. Abbeel, “Guided cost learning: Deep inverse optimal control via policy optimization,” in International Conference on Machine Learning . PMLR, 2016, pp. 49–58
2016
Earlier work this paper cites.
F. Johansson, U. Shalit, and D. Sontag, “Learning representations for counterfactual inference,” in International Conference on Machine Learning . PMLR, 2016, pp. 3020–3029
2016
Earlier work this paper cites.
H. Van Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double q-learning,” in Proceedings of the AAAI conference on artificial intelligence , vol. 30, no. 1, 2016
2016
Earlier work this paper cites.
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning,” International Conference on Learning Representations (ICLR) , 2016
2016
Earlier work this paper cites.
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot et al. , “Mastering the game of go with deep neural networks and tree search,” nature , vol. 529, no. 7587, pp. 484–489, 2016
2016
Earlier work this paper cites.
J. Zhang and E. Bareinboim, “Markov decision processes with unobserved confounders: A causal approach,” Technical report, Technical Report R-23, Purdue AI Lab, Tech. Rep., 2016
2016
Earlier work this paper cites.
F. Lattimore, T. Lattimore, and M. D. Reid, “Causal bandits: Learning good interventions via causal inference,” Advances in Neural Information Processing Systems , vol. 29, 2016
2016
Earlier work this paper cites.
G. Katz, D.-W. Huang, R. Gentili, and J. Reggia, “Imitation learning as cause-effect reasoning,” in International Conference on Artificial General Intelligence . Springer, 2016, pp. 64–73
2016
Earlier work this paper cites.
A. E. Johnson, T. J. Pollard, L. Shen, L.-w. H. Lehman, M. Feng, M. Ghassemi, B. Moody, P. Szolovits, L. Anthony Celi, and R. G. Mark, “Mimic-iii, a freely accessible critical care database,” Scientific data , vol. 3, no. 1, pp. 1–9, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
P. Jonas, J. Dominik, and S. Bernhard, Elements of Causal Inference: foundations and learning algorithms . MIT Press, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton et al. , “Mastering the game of go without human knowledge,” nature , vol. 550, no. 7676, pp. 354–359, 2017
2017
Earlier work this paper cites.
A. Forney, J. Pearl, and E. Bareinboim, “Counterfactual data-fusion for online reinforcement learners,” in International Conference on Machine Learning . PMLR, 2017, pp. 1156–1164
2017
Earlier work this paper cites.
R. Sen, K. Shanmugam, A. G. Dimakis, and S. Shakkottai, “Identifying best interventions through online importance sampling,” in International Conference on Machine Learning . PMLR, 2017, pp. 3057–3066
2017
Earlier work this paper cites.
J. Zhang and E. Bareinboim, “Transfer learning in multi-armed bandit: a causal approach,” in Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence (IJCAI) , 2017, pp. 1340–1346
2017
Earlier work this paper cites.
Z. Zhou, M. Bloem, and N. Bambos, “Infinite time horizon maximum causal entropy inverse reinforcement learning,” IEEE Transactions on Automatic Control , vol. 63, no. 9, pp. 2787–2802, 2017
2017
Earlier work this paper cites.
A. Hussein, M. M. Gaber, E. Elyan, and C. Jayne, “Imitation learning: A survey of learning methods,” ACM Computing Surveys (CSUR) , vol. 50, no. 2, pp. 1–35, 2017
2017
Earlier work this paper cites.
G. E. Katz, “A cognitive robotic imitation learning system based on cause-effect reasoning,” Ph.D. dissertation, University of Maryland, College Park, 2017
2017
Earlier work this paper cites.
G. Katz, D.-W. Huang, T. Hauge, R. Gentili, and J. Reggia, “A novel parsimonious cause-effect reasoning algorithm for robot imitation and plan recognition,” IEEE Transactions on Cognitive and Developmental Systems , vol. 10, no. 2, pp. 177–193, 2017
2017
Earlier work this paper cites.
K. Kansky, T. Silver, D. A. Mély, M. Eldawy, M. Lázaro-Gredilla, X. Lou, N. Dorfman, S. Sidor, S. Phoenix, and D. George, “Schema networks: Zero-shot transfer with a generative causal model of intuitive physics,” in International Conference on Machine Learning . PMLR, 2017, pp. 1809–1818
2017
Earlier work this paper cites.
Y. Li, “Deep reinforcement learning: An overview,” arXiv preprint arXiv:1701.07274 , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Pritzel, B. Uria, S. Srinivasan, A. P. Badia, O. Vinyals, D. Hassabis, D. Wierstra, and C. Blundell, “Neural episodic control,” in International Conference on Machine Learning . PMLR, 2017, pp. 2827–2836
2017
Earlier work this paper cites.
A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V. Koltun, “CARLA: An open urban driving simulator,” in Proceedings of the 1st Annual Conference on Robot Learning , 2017, pp. 1–16
2017
Cited alongside, same era.
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction . MIT press, 2018
2018
Cited alongside, same era.
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger, “Deep reinforcement learning that matters,” in Proceedings of the AAAI conference on artificial intelligence , vol. 32, no. 1, 2018
2018
Cited alongside, same era.
C. Lu, “Introduction to causal reinforcement learning,” Blogpost at causallu.com , 2018
2018
Cited alongside, same era.
J. Pearl and D. Mackenzie, The book of why: the new science of cause and effect . Basic books, 2018
2018
Cited alongside, same era.
Y. Sun, F. Zhuang, H. Zhu, Q. He, and H. Xiong, “Cost-effective and interpretable job skill recommendation with deep reinforcement learning,” in Proceedings of the Web Conference 2021 , 2021, pp. 3827–3838
2021
Later among the works it cites.
C. Yu, J. Liu, S. Nemati, and G. Yin, “Reinforcement learning in healthcare: A survey,” ACM Computing Surveys (CSUR) , vol. 55, no. 1, pp. 1–36, 2021
2021
Later among the works it cites.
S. J. Grimbly, J. Shock, and A. Pretorius, “Causal multi-agent reinforcement learning: Review and open problems,” in Cooperative AI Workshop in Advances in Neural Information Processing Systems , 2021
2021
Later among the works it cites.
L. Yao, Z. Chu, S. Li, Y. Li, J. Gao, and A. Zhang, “A survey on causal inference,” ACM Transactions on Knowledge Discovery from Data (TKDD) , vol. 15, no. 5, pp. 1–46, 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Huang, K. Zhang, Y. Lin, B. Schölkopf, and C. Glymour, “Generalized score functions for causal discovery,” in Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . ACM, 2018, pp. 1551–1560
2018
Cited alongside, same era.
D. Malinsky and P. Spirtes, “Causal structure learning from multivariate time series in settings with unmeasured confounding,” in Proceedings of 2018 ACM SIGKDD workshop on causal discovery . PMLR, 2018, pp. 23–47
2018
Cited alongside, same era.
W. Miao, Z. Geng, and E. J. Tchetgen Tchetgen, “Identifying causal effects with proxy variables of an unmeasured confounder,” Biometrika , vol. 105, no. 4, pp. 987–993, 2018
2018
Cited alongside, same era.
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” in International Conference on Machine Learning . PMLR, 2018, pp. 1861–1870
2018
Cited alongside, same era.
A. Nagabandi, G. Kahn, R. S. Fearing, and S. Levine, “Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning,” in 2018 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2018, pp. 7559–7566
2018
Cited alongside, same era.
K. Chua, R. Calandra, R. McAllister, and S. Levine, “Deep reinforcement learning in a handful of trials using probabilistic dynamics models,” Advances in Neural Information Processing Systems , vol. 31, 2018
2018
Cited alongside, same era.
T. Kurutach, I. Clavera, Y. Duan, A. Tamar, and P. Abbeel, “Model-ensemble trust-region policy optimization,” in International Conference on Learning Representations , 2018
2018
Cited alongside, same era.
X. Nie and S. Wager, “Quasi-oracle estimation of heterogeneous treatment effects,” Biometrika , vol. 108, no. 2, pp. 299–319, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
J. Li, Y. Luo, and X. Zhang, “Causal reinforcement learning: An instrumental variable approach,” Available at SSRN 3792824 , 2021
2021
Later among the works it cites.
L. Xu, H. Kanagawa, and A. Gretton, “Deep proxy causal learning and its application to confounded bandit policy evaluation,” in Advances in Neural Information Processing Systems , vol. 34, 2021, pp. 26 264–26 275
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
D. Zhu, L. E. Li, and M. Elhoseiny, “Causaldyna: Improving generalization of dyna-style reinforcement learning via counterfactual-based data augmentation,” 2021
2021
Later among the works it cites.
L. Wang, Z. Yang, and Z. Wang, “Provably efficient causal reinforcement learning with confounded observational data,” Advances in Neural Information Processing Systems , vol. 34, pp. 21 164–21 175, 2021
2021
Later among the works it cites.
D. A. Bruns-Smith, “Model-free and model-based policy evaluation when causality is uncertain,” in International Conference on Machine Learning . PMLR, 2021, pp. 1116–1126
2021
Later among the works it cites.
2021
Later among the works it cites.
Y. Li, H. Xie, Y. Lin, and J. C. Lui, “Unifying offline causal inference and online bandit learning for data driven decision,” in Proceedings of the Web Conference 2021 , 2021, pp. 2291–2303
2021
Later among the works it cites.
——, “Minimax-optimal policy learning under unobserved confounding,” Management Science , vol. 67, no. 5, pp. 2870–2890, 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
D. Kumor, J. Zhang, and E. Bareinboim, “Sequential causal imitation learning with unobserved confounders,” Advances in Neural Information Processing Systems , vol. 34, pp. 14 669–14 680, 2021
2021
Later among the works it cites.
I. Feliciano-Avelino, A. Méndez-Molina, E. F. Morales, and L. E. Sucar, “Causal based action selection policy for reinforcement learning,” in Mexican International Conference on Artificial Intelligence . Springer, 2021, pp. 213–227
2021
Later among the works it cites.
Y. Nair and N. Jiang, “A spectral approach to off-policy evaluation for pomdps,” The 38th International Conference on Machine Learning , 2021
2021
Later among the works it cites.
T. Herlau and R. Larsen, “Reinforcement learning of causal variables using mediation analysis,” in 36th AAAI Conference on Artificial Intelligence . Association for the Advancement of Artificial Intelligence, 2021
2021
Later among the works it cites.
A. Zhang, R. T. McAllister, R. Calandra, Y. Gal, and S. Levine, “Learning invariant representations for reinforcement learning without reconstruction,” in International Conference on Learning Representations , 2021
2021
Later among the works it cites.
M. Seitzer, B. Schölkopf, and G. Martius, “Causal influence detection for improving efficiency in reinforcement learning,” Advances in Neural Information Processing Systems , vol. 34, pp. 22 905–22 918, 2021
2021
Later among the works it cites.
S. A. Sontakke, A. Mehrjou, L. Itti, and B. Schölkopf, “Causal curiosity: Rl agents discovering self-supervised experiments for causal representation learning,” in International Conference on Machine Learning . PMLR, 2021, pp. 9848–9858
2021
Later among the works it cites.
A. Sonar, V. Pacelli, and A. Majumdar, “Invariant policy optimization: Towards stronger generalization in reinforcement learning,” in Learning for Dynamics and Control . PMLR, 2021, pp. 21–33
2021
Later among the works it cites.
Y. Lu, A. Meisami, and A. Tewari, “Causal bandits with unknown graph structure,” Advances in Neural Information Processing Systems , vol. 34, 2021
2021
Later among the works it cites.
S. Saengkyongam, N. Thams, J. Peters, and N. Pfister, “Invariant policy learning: A causal perspective,” Workshop on Reinforcement Learning Theory at the 38th International Conference on Machine Learning , 2021
2021
Later among the works it cites.
G. Tennenholtz, U. Shalit, S. Mannor, and Y. Efroni, “Bandits with partially observable confounded data,” in Uncertainty in Artificial Intelligence . PMLR, 2021, pp. 430–439
2021
Later among the works it cites.
I. Bica, D. Jarrett, A. Hüyük, and M. van der Schaar, “Learning” what-if” explanations for sequential decision-making,” International Conference on Learning Representations , 2021
2021
Later among the works it cites.
I. Bica, D. Jarrett, and M. van der Schaar, “Invariant causal imitation learning for generalizable policies,” Advances in Neural Information Processing Systems , vol. 34, 2021
2021
Later among the works it cites.
C. Wen, J. Lin, J. Qian, Y. Gao, and D. Jayaraman, “Keyframe-focused visual imitation learning,” in International Conference on Machine Learning . PMLR, 2021, pp. 11 123–11 133
2021
Later among the works it cites.
A. Méndez-Molina, “Combining reinforcement learning and causal models for robotics applications,” in Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI-21 , 8 2021, pp. 4905–4906. [Online]. Available: https://doi.org/10.24963/ijcai.2021/684
2021
Later among the works it cites.
2021
Later among the works it cites.
V. Luczkow, Structural Causal Models for Reinforcement Learning . McGill University (Canada), 2021
2021
Later among the works it cites.
M. Tomar, A. Zhang, R. Calandra, M. E. Taylor, and J. Pineau, “Model-invariant state abstractions for model-based reinforcement learning,” in Self-Supervision for Reinforcement Learning Workshop - ICLR 2021 , 2021
2021
Later among the works it cites.
C. Lyle, A. Zhang, M. Jiang, J. Pineau, and Y. Gal, “Resolving causal confusion in reinforcement learning via robust exploration,” in Self-Supervision for Reinforcement Learning Workshop of ICLR , 2021
2021
Later among the works it cites.
Y. Sun, K. Zhang, and C. Sun, “Model-based transfer reinforcement learning based on graphical model representations,” IEEE Transactions on Neural Networks and Learning Systems , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
——, “Inferring time-delayed causal relations in pomdps from the principle of independence of cause and mechanism,” in Proceedings of the 30th International Joint Conference on Artificial Intelligence (IJCAI) , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
T. E. Lee, J. A. Zhao, A. S. Sawhney, S. Girdhar, and O. Kroemer, “Causal reasoning in simulation for structure and transfer learning of robot manipulation policies,” in 2021 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2021, pp. 4776–4782
2021
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
Y. Chen, L. Xu, C. Gulcehre, T. Le Paine, A. Gretton, N. de Freitas, and A. Doucet, “On instrumental variable regression for deep offline policy evaluation,” Journal of Machine Learning Research , vol. 23, no. 302, pp. 1–40, 2022
2022
Later among the works it cites.
Z.-M. Zhu, S. Jiang, Y.-R. Liu, Y. Yu, and K. Zhang, “Invariant action effect model for reinforcement learning,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 8, 2022, pp. 9260–9268
2022
Later among the works it cites.
2022
Later among the works it cites.
J. Guo, M. Gong, and D. Tao, “A relational intervention approach for unsupervised dynamics generalization in model-based reinforcement learning,” in International Conference on Learning Representations , 2022
2022
Later among the works it cites.
T. W. Killian, M. Ghassemi, and S. Joshi, “Counterfactually guided policy transfer in clinical settings,” in Conference on Health, Inference, and Learning . PMLR, 2022, pp. 5–31
2022
Later among the works it cites.
2022
Later among the works it cites.
C. Shi, M. Uehara, J. Huang, and N. Jiang, “A minimax learning approach to off-policy evaluation in confounded partially observable markov decision processes,” in International Conference on Machine Learning . PMLR, 2022, pp. 20 057–20 094
2022
Later among the works it cites.
C. Subramanian and B. Ravindran, “Causal contextual bandits with targeted interventions,” in International Conference on Learning Representations , 2022
2022
Later among the works it cites.
——, “Online reinforcement learning for mixed policy scopes,” Columbia CausalAI Laboratory, Technical Report (R-84) , 2022
2022
Later among the works it cites.
G. Tennenholtz, A. Hallak, G. Dalal, S. Mannor, G. Chechik, and U. Shalit, “On covariate shift of latent confounders in imitation and reinforcement learning,” in International Conference on Learning Representations , 2022
2022
Later among the works it cites.
G. Swamy, S. Choudhury, J. A. Bagnell, and Z. S. Wu, “Causal imitation learning under temporally correlated noise,” in International Conference on Machine Learning . PMLR, 2022, pp. 20 877–20 890
2022
Later among the works it cites.
Y. Lu, A. Meisami, and A. Tewari, “Efficient reinforcement learning with prior causal knowledge,” in First Conference on Causal Learning and Reasoning , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
S. Peng, X. Hu, R. Zhang, K. Tang, J. Guo, Q. Yi, R. Chen, X. Zhang, Z. Du, L. Li, Q. Guo, and Y. Chen, “Causality-driven hierarchical structure discovery for reinforcement learning,” in Advances in Neural Information Processing Systems , A. H. Oh, A. Agarwal, D. Belgrave, and K. Cho, Eds., 2022
2022
Later among the works it cites.
Z. Wang, X. Xiao, Z. Xu, Y. Zhu, and P. Stone, “Causal dynamics learning for task-independent state abstraction,” in International Conference on Machine Learning . PMLR, 2022, pp. 23 151–23 180
2022
Later among the works it cites.
B. Huang, C. Lu, L. Leqi, J. M. Hernández-Lobato, C. Glymour, B. Schölkopf, and K. Zhang, “Action-sufficient state representation learning for control with structural constraints,” in International Conference on Machine Learning . International Machine Learning Society, 2022
2022
Later among the works it cites.
B. Huang, F. Feng, C. Lu, S. Magliacane, and K. Zhang, “Adarl: What, where, and how to adapt in transfer reinforcement learning,” International Conference on Learning Representations , 2022
2022
Later among the works it cites.
C.-C. Chuang, D. Yang, C. Wen, and Y. Gao, “Resolving copycat problems in visual imitation learning via residual action prediction,” in 17th European Conference on Computer Vision . Springer, 2022, pp. 392–409
2022
Later among the works it cites.
C. Wen, J. Qian, J. Lin, J. Teng, D. Jayaraman, and Y. Gao, “Fighting fire with fire: avoiding dnn shortcuts through priming,” in International Conference on Machine Learning . PMLR, 2022, pp. 23 723–23 750
2022
Later among the works it cites.
W. Ding, H. Lin, B. Li, and D. Zhao, “Generalizing goal-conditioned reinforcement learning with variational causal reasoning,” in Advances in Neural Information Processing Systems , 2022
2022
Later among the works it cites.
C. Lu, J. M. Hernández-Lobato, and B. Schölkopf, “Invariant causal representation learning for generalization in imitation and reinforcement learning,” in ICLR2022 Workshop on the Elements of Reasoning: Objects, Structure and Causality , 2022
2022
Later among the works it cites.
2022
Later among the works it cites.
Wikipedia contributors, “Cognition,” 2022, accessed 18-November-2022. [Online]. Available: https://en.wikipedia.org/w/index.php?title=Cognition&oldid=1118273677
2022
Later among the works it cites.
V. Nair, V. Patil, and G. Sinha, “Budgeted and non-budgeted causal bandits,” in International Conference on Artificial Intelligence and Statistics . PMLR, 2021, pp. 2017–2025
2025
Closest in time.