Fetching the paper…
Reading the bibliography…
Dyna is a fundamental approach to model-based reinforcement learning (MBRL) that interleaves planning, acting, and learning in an online setting.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S. Sutton · 1990
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
Long-Ji Lin · 1992
Earlier work this paper cites.
Prioritized sweeping: Reinforcement learning with less data and less time
Andrew W. Moore and Christopher G. Atkeson · 1993
Earlier work this paper cites.
Efficient learning and planning within the Dyna framework
Jing Peng and Ronald J. Williams · 1993
Earlier work this paper cites.
On-line Q-learning using connectionist systems
G. A. Rummery and M. Niranjan · 1994
Earlier work this paper cites.
Generalization in reinforcement learning: Successful examples using sparse coarse coding
R. S. Sutton · 1996
Earlier work this paper cites.
Agnostic system identification for model-based reinforcement learning
Stephane Ross and Drew Bagnell · 2012
Earlier work this paper cites.
Lecture 6.5 – RMSProp: Divde the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Earlier work this paper cites.
Bayesian learning of recursively factored environments
Marc G. Bellemare, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Skip context tree switching
Marc G. Bellemare, Joel Veness, and Erik Talvitie · 2014
Cited alongside, same era.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Cited alongside, same era.
Model regularization for stable sample rollouts
Erik Talvitie · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Action-Conditional Video Prediction using Deep Networks in Atari Games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee, Richard L Lewis, and Satinder Singh · 2015
Cited alongside, same era.
Recurrent environment simulators
Silvia Chiappa, Sébastien Racaniere, Daan Wierstra, and Shakir Mohamed · 2017
Later among the works it cites.
Uncertainty-driven imagination for continuous deep reinforcement learning
Gabriel Kalweit and Joschka Boedecker · 2017
Later among the works it cites.
A deep learning approach for joint video frame and reward prediction in atari games
Felix Leibfried, Nate Kushman, and Katja Hofmann · 2017
Later among the works it cites.
Marlos C. Machado, Marc G. Bellemare, Erik Talvitie, Joel Veness, Matthew J. Hausknecht, and Michael Bowling · 2017
Later among the works it cites.
Value prediction network
Junhyuk Oh, Satinder Singh, and Honglak Lee · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Action-conditional video prediction using deep networks in atari games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee, Richard L Lewis, and Satinder Singh · 2015
Cited alongside, same era.
Continuous deep Q-learning with model-based acceleration
Shixiang Gu, Timothy Lillicrap, Ilya Sutskever, and Sergey Levine · 2016
Cited alongside, same era.
State of the art control of Atari games using shallow reinforcement learning
Yitao Liang, Marlos C. Machado, Erik Talvitie, and Michael Bowling · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Value iteration networks
Aviv Tamar, Yi Wu, Garrett Thomas, Sergey Levine, and Pieter Abbeel · 2016
Cited alongside, same era.
Theophane Weber, Sébastien Racanière, David Reichert, Lars Buesing, Arthur Guez, Danilo Jimenez Rezende, Adrià Puigdomènech Badia, Oriol Vinyals, Nicolas Heess, Yujia Li, Razvan Pascanu, Peter Battaglia, Demis Hassabis, David Silver, and Daan Wierstra · 2017
Later among the works it cites.
Recall traces: Backtracking models for efficient reinforcement learning
Anirudh Goyal, Philemon Brakel, William Fedus, Timothy P. Lillicrap, Sergey Levine, Hugo Larochelle, and Yoshua Bengio · 2018
Closest in time.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2018
Closest in time.
Organizing experience: a deeper look at replay mechanisms for sample-based planning in continuous state domains
Yangchen Pan, Muhammad Zaheer, Adam White, Andrew Patterson, and Martha White · 2018
Closest in time.
Deep Dyna-Q: Integrating planning for task-completion dialogue policy learning
Baolin Peng, Xiujun Li, Jianfeng Gao, Jingjing Liu, and Kam-Fai Wong · 2018
Closest in time.