Fetching the paper…
Reading the bibliography…
Model-based methods have recently shown promising for offline reinforcement learning (RL), aiming to learn good policies from historical data without interacting with the environment.
A markovian decision process
Richard Bellman · 1957
Earlier work this paper cites.
Causation, Prediction, and Search
Peter Spirtes, Clark N Glymour, and Richard Scheines · 2000
Earlier work this paper cites.
Ridge regression: Biased estimation for nonorthogonal problems
Arthur E. Hoerl and Robert W. Kennard · 2000
Earlier work this paper cites.
Models, reasoning and inference
Judea Pearl et al · 2000
Earlier work this paper cites.
Distribution-free learning of bayesian network structure in continuous domains
Dimitris Margaritis · 2005
Earlier work this paper cites.
A kernel-based causal learning algorithm
Xiaohai Sun, Dominik Janzing, Bernhard Schölkopf, and Kenji Fukumizu · 2007
Earlier work this paper cites.
Probabilistic Graphical Models - Principles and Techniques
Daphne Koller and Nir Friedman · 2009
Earlier work this paper cites.
Kernel-based conditional independence test and application in causal discovery
Kun Zhang, Jonas Peters, Dominik Janzing, and Bernhard Schölkopf · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Information theoretic MPC for model-based reinforcement learning
Grady Williams, Nolan Wagener, Brian Goldfain, Paul Drews, James M. Rehg, Byron Boots, and Evangelos A. Theodorou · 2017
Earlier work this paper cites.
BDD100K: A diverse driving video database with scalable annotation tooling
Fisher Yu, Wenqi Xian, Yingying Chen, Fangchen Liu, Mike Liao, Vashisht Madhavan, and Trevor Darrell · 2018
Cited alongside, same era.
Recurrent world models facilitate policy evolution
David Ha and Jürgen Schmidhuber · 2018
Cited alongside, same era.
Human causal transfer: Challenges for deep reinforcement learning
Mark Edmonds, James Kubricht, Colin Summers, Yixin Zhu, Brandon Rothrock, Song-Chun Zhu, and Hongjing Lu · 2018
Cited alongside, same era.
Building machines that learn and think like people
Josh Tenenbaum · 2018
Cited alongside, same era.
Model-ensemble trust-region policy optimization
Thanard Kurutach, Ignasi Clavera, Yan Duan, Aviv Tamar, and Pieter Abbeel · 2018
Cited alongside, same era.
Virtual-Taobao: Virtualizing real-world online retail environment for reinforcement learning
Jing-Cheng Shi, Yang Yu, Qing Da, Shi-Yong Chen, and An-Xiang Zeng · 2019
Solving rubik’s cube with a robot hand
OpenAI, Ilge Akkaya, Marcin Andrychowicz, Maciek Chociej, Mateusz Litwin, Bob McGrew, Arthur Petron, Alex Paino, Matthias Plappert, Glenn Powell, Raphael Ribas, Jonas Schneider, Nikolas Tezak, Jerry Tworek, Peter Welinder, Lilian Weng, Qiming Yuan, Wojciech Zaremba, and Lei Zhang · 2019
Later among the works it cites.
MOPO: model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Y. Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Later among the works it cites.
Error bounds of imitating policies and environments
Tian Xu, Ziniu Li, and Yang Yu · 2020
Later among the works it cites.
A meta-transfer objective for learning to disentangle causal mechanisms
Yoshua Bengio, Tristan Deleu, Nasim Rahaman, Nan Rosemary Ke, Sébastien Lachapelle, Olexa Bilaniuk, Anirudh Goyal, and Christopher J. Pal · 2020
Later among the works it cites.
Morel: Model-based offline reinforcement learning
Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, and Thorsten Joachims · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Guidelines for reinforcement learning in healthcare
O. Gottesman, F. Johansson, M. Komorowski, A. Faisal, David Sontag, Finale Doshi-Velez, and L. A. Celi · 2019
Cited alongside, same era.
Causal confusion in imitation learning
Pim de Haan, Dinesh Jayaraman, and Sergey Levine · 2019
Cited alongside, same era.
Learning neural causal models from unknown interventions
Nan Rosemary Ke, Olexa Bilaniuk, Anirudh Goyal, Stefan Bauer, Hugo Larochelle, Chris Pal, and Yoshua Bengio · 2019
Cited alongside, same era.
Partially observable environment estimation with uplift inference for reinforcement learning based recommendation
Wenjie Shang, Qingyang Li, Zhiwei Qin, Yang Yu, Yiping Meng, and Jieping Ye · 2021
Later among the works it cites.
Offline model-based adaptable policy learning
Xiong-Hui Chen, Yang Yu, Qingyang Li, Fan-Ming Luo, Zhiwei Tony Qin, Shang Wenjie, and Jieping Ye · 2021
Later among the works it cites.
Adapting environment sudden changes by learning context sensitive policy
Fan-Ming Luo, Shengyi Jiang, Yang Yu, Zongzhang Zhang, and Yi-Feng Zhang · 2022
Closest in time.
Invariant action effect model for reinforcement learning
Zheng-Mao Zhu, Shengyi Jiang, Yu-Ren Liu, Yang Yu, and Kun Zhang · 2022
Closest in time.