Fetching the paper…
Reading the bibliography…
As a pivotal component to attaining generalizable solutions in human intelligence, reasoning provides great potential for reinforcement learning (RL) agents' generalization towards varied goals by summarizing part-to-whole arguments and discovering cause-and-effect relations.
Stochastic dynamic programming with factored representations
Craig Boutilier, Richard Dearden, and Moisés Goldszmidt · 2000
Earlier work this paper cites.
Causation, prediction, and search
Peter Spirtes, Clark N Glymour, Richard Scheines, and David Heckerman · 2000
Earlier work this paper cites.
Optimal structure identification with greedy search
David Maxwell Chickering · 2002
Earlier work this paper cites.
Robust constrained model predictive control
Arthur George Richards · 2005
Earlier work this paper cites.
Discovering symbolic models from deep learning with inductive biases (2020)
Miles Cranmer, Alvaro Sanchez-Gonzalez, Peter Battaglia, Rui Xu, Kyle Cranmer, David Spergel, and Shirley Ho · 2006
Earlier work this paper cites.
The max-min hill-climbing bayesian network structure learning algorithm
Ioannis Tsamardinos, Laura E Brown, and Constantin F Aliferis · 2006
Earlier work this paper cites.
No-regret reductions for imitation learning and structured prediction
Stéphane Ross, Geoffrey J Gordon, and J Andrew Bagnell · 2011
Earlier work this paper cites.
Kernel-based conditional independence test and application in causal discovery
Kun Zhang, Jonas Peters, Dominik Janzing, and Bernhard Schölkopf · 2012
Earlier work this paper cites.
Characterization and greedy learning of interventional markov equivalence classes of directed acyclic graphs
Alain Hauser and Peter Bühlmann · 2012
Earlier work this paper cites.
The bayesian information criterion: background, derivation, and applications
Andrew A Neath and Joseph E Cavanaugh · 2012
Earlier work this paper cites.
Model predictive control
Eduardo F Camacho and Carlos Bordons Alba · 2013
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Universal value function approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Earlier work this paper cites.
Towards deep symbolic reinforcement learning
Marta Garnelo, Kai Arulkumaran, and Murray Shanahan · 2016
Earlier work this paper cites.
Reinforcement learning and causal models
Samuel J Gershman · 2017
Earlier work this paper cites.
Elements of causal inference: foundations and learning algorithms
Jonas Peters, Dominik Janzing, and Bernhard Schölkopf · 2017
Earlier work this paper cites.
Hindsight experience replay
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba · 2017
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel · 2017
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Earlier work this paper cites.
State abstractions for lifelong reinforcement learning
David Abel, Dilip Arumugam, Lucas Lehnert, and Michael Littman · 2018
Earlier work this paper cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Sergey Levine · 2018
Earlier work this paper cites.
Maximum a posteriori policy optimisation
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Remi Munos, Nicolas Heess, and Martin Riedmiller · 2018
Earlier work this paper cites.
Fast conditional independence test for vector variables with large sample sizes
Krzysztof Chalupka, Pietro Perona, and Frederick Eberhardt · 2018
Earlier work this paper cites.
Testing conditional independence of discrete distributions
Clément L Canonne, Ilias Diakonikolas, Daniel M Kane, and Alistair Stewart · 2018
Earlier work this paper cites.
Minimalistic gridworld environment for openai gym
Maxime Chevalier-Boisvert, Lucas Willems, and Suman Pal · 2018
Earlier work this paper cites.
An environment for autonomous driving decision-making
Edouard Leurent · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine · 2018
Earlier work this paper cites.
Modeling relational data with graph convolutional networks
Michael Schlichtkrull, Thomas N Kipf, Peter Bloem, Rianne van den Berg, Ivan Titov, and Max Welling · 2018
Earlier work this paper cites.
Automatic goal generation for reinforcement learning agents
Carlos Florensa, David Held, Xinyang Geng, and Pieter Abbeel · 2018
Earlier work this paper cites.
Experimental design for cost-aware learning of causal graphs
Erik Lindgren, Murat Kocaoglu, Alexandros G Dimakis, and Sriram Vishwanath · 2018
Earlier work this paper cites.
Characterizing and learning equivalence classes of causal dags under interventions
Karren D. Yang, Abigail Katoff, and Caroline Uhler · 2018
Earlier work this paper cites.
Program guided agent
Shao-Hua Sun, Te-Lin Wu, and Joseph J Lim · 2019
Cited alongside, same era.
Causal induction from visual observations for goal directed tasks
Suraj Nair, Yuke Zhu, Silvio Savarese, and Li Fei-Fei · 2019
Cited alongside, same era.
An inference perspective on model-based reinforcement learning
Joseph Marino and Yisong Yue · 2019
Cited alongside, same era.
Benchmarking model-based reinforcement learning
Tingwu Wang, Xuchan Bao, Ignasi Clavera, Jerrick Hoang, Yeming Wen, Eric Langlois, Shunshi Zhang, Guodong Zhang, Pieter Abbeel, and Jimmy Ba · 2019
Cited alongside, same era.
Compile: Compositional imitation learning and execution
Thomas Kipf, Yujia Li, Hanjun Dai, Vinicius Zambaldi, Alvaro Sanchez-Gonzalez, Edward Grefenstette, Pushmeet Kohli, and Peter Battaglia · 2019
Cited alongside, same era.
Explainable reinforcement learning through a causal lens
Prashan Madumal, Tim Miller, Liz Sonenberg, and Frank Vetere · 2020
Later among the works it cites.
Causal discovery in physical systems from videos
Yunzhu Li, Antonio Torralba, Anima Anandkumar, Dieter Fox, and Animesh Garg · 2020
Later among the works it cites.
Efficient intervention design for causal discovery with latents
Raghavendra Addanki, Shiva Kasiviswanathan, Andrew Mcgregor, and Cameron Musco · 2020
Later among the works it cites.
A bayesian-symbolic approach to reasoning and learning in intuitive physics
Kai Xu, Akash Srivastava, Dan Gutfreund, Felix Sosa, Tomer Ullman, Josh Tenenbaum, and Charles Sutton · 2021
Later among the works it cites.
Learning task decomposition with ordered memory policy network
Yuchen Lu, Yikang Shen, Siyuan Zhou, Aaron Courville, Joshua B Tenenbaum, and Chuang Gan · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Regression planning networks
Danfei Xu, Roberto Martín-Martín, De-An Huang, Yuke Zhu, Silvio Savarese, and Li F Fei-Fei · 2019
Cited alongside, same era.
Neural task graphs: Generalizing to unseen tasks from a single video demonstration
De-An Huang, Suraj Nair, Danfei Xu, Yuke Zhu, Animesh Garg, Li Fei-Fei, Silvio Savarese, and Juan Carlos Niebles · 2019
Cited alongside, same era.
Network randomization: A simple technique for generalization in deep reinforcement learning
Kimin Lee, Kibok Lee, Jinwoo Shin, and Honglak Lee · 2019
Cited alongside, same era.
Structured domain randomization: Bridging the reality gap by context-aware synthetic data
Aayush Prakash, Shaad Boochoon, Mark Brophy, David Acuna, Eric Cameracci, Gavriel State, Omer Shapira, and Stan Birchfield · 2019
Cited alongside, same era.
Keeping your distance: Solving sparse reward tasks using self-balancing shaped rewards
Alexander Trott, Stephan Zheng, Caiming Xiong, and Richard Socher · 2019
Cited alongside, same era.
Exploration via hindsight goal generation
Zhizhou Ren, Kefan Dong, Yuan Zhou, Qiang Liu, and Jian Peng · 2019
Cited alongside, same era.
Curriculum-guided hindsight experience replay
Meng Fang, Tianyi Zhou, Yali Du, Lei Han, and Zhengyou Zhang · 2019
Cited alongside, same era.
Program synthesis guided reinforcement learning for partially observed environments
Yichen Yang, Jeevana Priya Inala, Osbert Bastani, Yewen Pu, Armando Solar-Lezama, and Martin Rinard · 2021
Later among the works it cites.
Discovering symbolic policies with deep reinforcement learning
Mikel Landajuela, Brenden K Petersen, Sookyung Kim, Claudio P Santiago, Ruben Glatt, Nathan Mundhenk, Jacob F Pettit, and Daniel Faissol · 2021
Later among the works it cites.
Proto: Program-guided transformer for program-guided tasks
Zelin Zhao, Karan Samel, Binghong Chen, et al · 2021
Later among the works it cites.
Model-invariant state abstractions for model-based reinforcement learning
Manan Tomar, Amy Zhang, Roberto Calandra, Matthew E Taylor, and Joelle Pineau · 2021
Later among the works it cites.
Causal curiosity: Rl agents discovering self-supervised experiments for causal representation learning
Sumedh A Sontakke, Arash Mehrjou, Laurent Itti, and Bernhard Schölkopf · 2021
Later among the works it cites.
Invariant causal imitation learning for generalizable policies
Ioana Bica, Daniel Jarrett, and Mihaela van der Schaar · 2021
Later among the works it cites.
Learning domain invariant representations in goal-conditioned block mdps
Beining Han, Chongyi Zheng, Harris Chan, Keiran Paster, Michael Zhang, and Jimmy Ba · 2021
Later among the works it cites.
Task-independent causal state abstraction
Zizhao Wang, Xuesu Xiao, Yuke Zhu, and Peter Stone · 2021
Later among the works it cites.
Causal influence detection for improving efficiency in reinforcement learning
Maximilian Seitzer, Bernhard Schölkopf, and Georg Martius · 2021
Later among the works it cites.
Causal reinforcement learning using observational and interventional data
Maxime Gasse, Damien Grasset, Guillaume Gaudron, and Pierre-Yves Oudeyer · 2021
Later among the works it cites.
Systematic evaluation of causal discovery in visual model based reinforcement learning
Nan Rosemary Ke, Aniket Didolkar, Sarthak Mittal, Anirudh Goyal, Guillaume Lajoie, Stefan Bauer, Danilo Rezende, Yoshua Bengio, Michael Mozer, and Christopher Pal · 2021
Later among the works it cites.
A language for counterfactual generative models
Zenna Tavares, James Koppel, Xin Zhang, Ria Das, and Armando Solar-Lezama · 2021
Later among the works it cites.
Stabilizing deep q-learning with convnets and vision transformers under data augmentation
Nicklas Hansen, Hao Su, and Xiaolong Wang · 2021
Later among the works it cites.
Generalization in reinforcement learning by soft data augmentation
Nicklas Hansen and Xiaolong Wang · 2021
Later among the works it cites.
Automatic data augmentation for generalization in reinforcement learning
Roberta Raileanu, Maxwell Goldstein, Denis Yarats, Ilya Kostrikov, and Rob Fergus · 2021
Later among the works it cites.
Hindsight expectation maximization for goal-conditioned reinforcement learning
Yunhao Tang and Alp Kucukelbir · 2021
Later among the works it cites.
Outcome-driven reinforcement learning via variational inference
Tim GJ Rudner, Vitchyr Pong, Rowan McAllister, Yarin Gal, and Sergey Levine · 2021
Later among the works it cites.
Causalaf: Causal autoregressive flow for goal-directed safety-critical scenes generation
Wenhao Ding, Haohong Lin, Bo Li, and Ding Zhao · 2021
Later among the works it cites.
Learning neural causal models with active interventions
Nino Scherrer, Olexa Bilaniuk, Yashas Annadani, Anirudh Goyal, Patrick Schwab, Bernhard Schölkopf, Michael C Mozer, Yoshua Bengio, Stefan Bauer, and Nan Rosemary Ke · 2021
Later among the works it cites.
Improving generalization with approximate factored value functions
Shagun Sodhani, Sergey Levine, and Amy Zhang · 2022
Closest in time.
A theory of abstraction in reinforcement learning
David Abel · 2022
Closest in time.
Abstraction for deep reinforcement learning
Murray Shanahan and Melanie Mitchell · 2022
Closest in time.
Goal-conditioned reinforcement learning: Problems and solutions
Minghuan Liu, Menghui Zhu, and Weinan Zhang · 2022
Closest in time.
Offline reinforcement learning with causal structured world models
Zheng-Mao Zhu, Xiong-Hui Chen, Hong-Long Tian, Kun Zhang, and Yang Yu · 2022
Closest in time.
Causal dynamics learning for task-independent state abstraction
Zizhao Wang, Xuesu Xiao, Zifan Xu, Yuke Zhu, and Peter Stone · 2022
Closest in time.
A survey on safety-critical scenario generation from methodological perspective
Wenhao Ding, Chejian Xu, Haohong Lin, Bo Li, and Ding Zhao · 2022
Closest in time.