Fetching the paper…
Reading the bibliography…
Deep model-based Reinforcement Learning (RL) has the potential to substantially improve the sample-efficiency of deep RL.
Dyna, an integrated architecture for learning, planning, and reacting
Richard S Sutton · 1991
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
Model-based influences on humans’ choices and striatal prediction errors
Nathaniel D Daw, Samuel J Gershman, Ben Seymour, Peter Dayan, and Raymond J Dolan · 2011
Earlier work this paper cites.
Model regularization for stable sample rollouts
Erik Talvitie · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Earlier work this paper cites.
A deeper look at planning as learning from replay
Harm Van Seijen and Rich Sutton · 2015
Cited alongside, same era.
True online temporal-difference learning
Harm Van Seijen, A Rupam Mahmood, Patrick M Pilarski, Marlos C Machado, and Richard S Sutton · 2016
Cited alongside, same era.
Imagination-augmented agents for deep reinforcement learning
Sébastien Racanière, Théophane Weber, David Reichert, Lars Buesing, Arthur Guez, Danilo Jimenez Rezende, Adria Puigdomenech Badia, Oriol Vinyals, Nicolas Heess, Yujia Li, et al · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Cited alongside, same era.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine · 2018
Cited alongside, same era.
Learning latent dynamics for planning from pixels
Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson · 2018
Later among the works it cites.
Dota 2 with large scale deep reinforcement learning
Christopher Berner, Greg Brockman, Brooke Chan, Vicki Cheung, Przemysław Dębiak, Christy Dennison, David Farhi, Quirin Fischer, Shariq Hashme, Chris Hesse, et al · 2019
Later among the works it cites.
Alphastar: Mastering the real-time strategy game StarCraft II, 2019
DeepMind · 2019
Later among the works it cites.
Dream to control: Learning behaviors by latent imagination
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi · 2019
Later among the works it cites.
Variational inference mpc for bayesian model-based reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gregory Farquhar, Tim Rocktaeschel, Maximilian Igl, and Shimon Whiteson · 2018
Cited alongside, same era.
Recurrent world models facilitate policy evolution
David Ha and Jürgen Schmidhuber · 2018
Cited alongside, same era.
URL https://github.com/Kaixhin/PlaNet
Kai Arulkumaran
Cited in the paper.
URL https://github.com/koulanurag/muzero-pytorch
Anurag Koul
Cited in the paper.
Masashi Okada and Tadahiro Taniguchi · 2019
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2019
Later among the works it cites.
Exploring model-based planning with policy networks
Tingwu Wang and Jimmy Ba · 2020
Closest in time.