Fetching the paper…
Reading the bibliography…
Deep reinforcement learning has shown remarkable success in the past few years.
Gradient theory of optimal flight paths
Henry J Kelley · 1960
Earlier work this paper cites.
Stabilizing state-feedback design via the moving horizon method
W Hi Kwon, AM Bruckstein, and T Kailath · 1983
Earlier work this paper cites.
Model predictive control: Theory and practice—a survey
Carlos E Garcia, David M Prett, and Manfred Morari · 1989
Earlier work this paper cites.
Learning from delayed rewards
Christopher JCH Watkins · 1989
Earlier work this paper cites.
An on-line algorithm for dynamic reinforcement learning and planning in reactive environments
Jürgen Schmidhuber · 1990
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S Sutton · 1990
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Richard S Sutton · 1991
Earlier work this paper cites.
Random decision forests
Tin Kam Ho · 1995
Earlier work this paper cites.
Reinforcement learning: A survey
Leslie Pack Kaelbling, Michael L Littman, and Andrew W Moore · 1996
Earlier work this paper cites.
The maxq method for hierarchical reinforcement learning
Thomas G Dietterich · 1998
Earlier work this paper cites.
Actor–critic algorithms
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
Deep Blue
Murray Campbell, A Joseph Hoane Jr, and Feng-hsiung Hsu · 2002
Earlier work this paper cites.
Robust constrained model predictive control
Arthur George Richards · 2005
Earlier work this paper cites.
Pattern recognition and machine learning
Christopher M Bishop · 2006
Earlier work this paper cites.
Efficient selectivity and backup operators in Monte-Carlo Tree Search
Rémi Coulom · 2006
Earlier work this paper cites.
A framework for reinforcement learning and planning
Thomas M Moerland, Joost Broekens, and Catholijn M Jonker · 2006
Earlier work this paper cites.
Model-based reinforcement learning: A survey
Thomas M Moerland, Joost Broekens, and Catholijn M Jonker · 2006
Earlier work this paper cites.
An application of reinforcement learning to aerobatic helicopter flight
Pieter Abbeel, Adam Coates, Morgan Quigley, and Andrew Y Ng · 2007
Earlier work this paper cites.
Efficient learning in cellular simultaneous recurrent neural networks—the case of maze navigation problem
Roman Ilin, Robert Kozma, and Paul J Werbos · 2007
Earlier work this paper cites.
Metalearning: Applications to data mining
Pavel Brazdil, Christophe Giraud Carrier, Carlos Soares, and Ricardo Vilalta · 2008
Earlier work this paper cites.
Dimensionality reduction: a comparative
Laurens Van Der Maaten, Eric Postma, Jaap Van den Herik, et al · 2009
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Marc Deisenroth and Carl E Rasmussen · 2011
Earlier work this paper cites.
Thinking, fast and slow
Daniel Kahneman · 2011
Earlier work this paper cites.
Multi-armed bandits with episode context
Christopher D Rosin · 2011
Earlier work this paper cites.
A survey of Monte Carlo Tree Search methods
Cameron B Browne, Edward Powley, Daniel Whitehouse, Simon M Lucas, Peter I Cowling, Philipp Rohlfshagen, Stephen Tavener, Diego Perez, Spyridon Samothrakis, and Simon Colton · 2012
Earlier work this paper cites.
Temporal-difference search in computer Go
David Silver, Richard S Sutton, and Martin Müller · 2012
Earlier work this paper cites.
Synthesis and stabilization of complex behaviors through online trajectory optimization
Yuval Tassa, Tom Erez, and Emanuel Todorov · 2012
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
The Arcade Learning Environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Dynamic programming
Richard Bellman · 2013
Earlier work this paper cites.
The cross-entropy method for optimization
Zdravko I Botev, Dirk P Kroese, Reuven Y Rubinstein, and Pierre L’Ecuyer · 2013
Earlier work this paper cites.
Sokoban.org, 2013
Yang Chao · 2013
Earlier work this paper cites.
A survey on policy search for robotics
Marc Peter Deisenroth, Gerhard Neumann, and Jan Peters · 2013
Earlier work this paper cites.
Auto-encoding variational Bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Jens Kober, J Andrew Bagnell, and Jan Peters · 2013
Earlier work this paper cites.
Guided policy search
Sergey Levine and Vladlen Koltun · 2013
Earlier work this paper cites.
Playing Atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
A survey of real-time strategy game AI research and competition in StarCraft
Santiago Ontanón, Gabriel Synnaeve, Alberto Uriarte, Florian Richoux, David Churchill, and Mike Preuss · 2013
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Cited alongside, same era.
Aircraft optimal terrain/threat-based trajectory planning and control
Reza Kamyar and Ehsan Taheri · 2014
Cited alongside, same era.
Learning neural network policies with guided policy search under unknown dynamics
Sergey Levine and Pieter Abbeel · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Cited alongside, same era.
Learning continuous control policies by stochastic value gradients
Nicolas Heess, Gregory Wayne, David Silver, Timothy Lillicrap, Tom Erez, and Yuval Tassa · 2015
Cited alongside, same era.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Cited alongside, same era.
TreeQN and ATreeC: Differentiable tree planning for deep reinforcement learning
Gregory Farquhar, Tim Rocktäschel, Maximilian Igl, and SA Whiteson · 2018
Later among the works it cites.
Model-based value estimation for efficient model-free reinforcement learning
Vladimir Feinberg, Alvin Wan, Ion Stoica, Michael I Jordan, Joseph E Gonzalez, and Sergey Levine · 2018
Later among the works it cites.
Learning to search with MCTSnets
Arthur Guez, Théophane Weber, Ioannis Antonoglou, Karen Simonyan, Oriol Vinyals, Daan Wierstra, Rémi Munos, and David Silver · 2018
Later among the works it cites.
Learning latent dynamics for planning from pixels
Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson · 2018
Later among the works it cites.
Model-based reinforcement learning https://medium.com/@jonathan_hui/rl-model-based-reinforcement-learning-3c2b6f0aa323
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Action-conditional video prediction using deep networks in Atari games
Junhyuk Oh, Xiaoxiao Guo, Honglak Lee, Richard L Lewis, and Satinder Singh · 2015
Cited alongside, same era.
Agnostic system identification for monte carlo planning
Erik Talvitie · 2015
Cited alongside, same era.
Convolutional LSTM network: A machine learning approach for precipitation nowcasting
SHI Xingjian, Zhourong Chen, Hao Wang, Dit-Yan Yeung, Wai-Kin Wong, and Wang-chun Woo · 2015
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Cited alongside, same era.
Jonathan Hui · 2018
Later among the works it cites.
Learn data science webpage., 2018
Satwik Kansal and Brendan Martin · 2018
Later among the works it cites.
A0c: Alpha zero in continuous action space
Thomas M Moerland, Joost Broekens, Aske Plaat, and Catholijn M Jonker · 2018
Later among the works it cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Anusha Nagabandi, Gregory Kahn, Ronald S Fearing, and Sergey Levine · 2018
Later among the works it cites.
Nantas Nardelli, Gabriel Synnaeve, Zeming Lin, Pushmeet Kohli, Philip HS Torr, and Nicolas Usunier · 2018
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis · 2018
Later among the works it cites.
Aravind Srinivas, Allan Jabri, Pieter Abbeel, Sergey Levine, and Chelsea Finn · 2018
Later among the works it cites.
Reinforcement learning, An Introduction, Second Edition
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, et al · 2018
Later among the works it cites.
Deep reinforcement learning for general video game ai
Ruben Rodriguez Torrado, Philip Bontrager, Julian Togelius, Jialin Liu, and Diego Perez-Liebana · 2018
Later among the works it cites.
Relational deep reinforcement learning
Vinicius Zambaldi, David Raposo, Adam Santoro, Victor Bapst, Yujia Li, Igor Babuschkin, Karl Tuyls, David Reichert, Timothy Lillicrap, Edward Lockhart, et al · 2018
Later among the works it cites.
Superhuman AI for multiplayer poker
Noam Brown and Tuomas Sandholm · 2019
Later among the works it cites.
Model-free reinforcement learning algorithms: A survey
Sinan Çalışır and Meltem Kurt Pehlivanoğlu · 2019
Later among the works it cites.
An investigation of model-free planning
Arthur Guez, Mehdi Mirza, Karol Gregor, Rishabh Kabra, Sébastien Racanière, Théophane Weber, David Raposo, Adam Santoro, Laurent Orseau, Tom Eccles, et al · 2019
Later among the works it cites.
Dream to control: Learning behaviors by latent imagination
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi · 2019
Later among the works it cites.
Analogues of mental simulation and imagination in deep learning
Jessica B Hamrick · 2019
Later among the works it cites.
When to trust your model: Model-based policy optimization
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine · 2019
Later among the works it cites.
Deep learning for video game playing
Niels Justesen, Philip Bontrager, Julian Togelius, and Sebastian Risi · 2019
Later among the works it cites.
Model-based reinforcement learning for Atari
Lukasz Kaiser, Mohammad Babaeizadeh, Piotr Milos, Blazej Osinski, Roy H Campbell, Konrad Czechowski, Dumitru Erhan, Chelsea Finn, Piotr Kozakowski, Sergey Levine, et al · 2019
Later among the works it cites.
An introduction to variational autoencoders
Diederik P Kingma and Max Welling · 2019
Later among the works it cites.
Value iteration networks on multiple levels of abstraction
Daniel Schleich, Tobias Klamt, and Sven Behnke · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Later among the works it cites.
Benchmarking model-based reinforcement learning
Tingwu Wang, Xuchan Bao, Ignasi Clavera, Jerrick Hoang, Yeming Wen, Eric Langlois, Shunshi Zhang, Guodong Zhang, Pieter Abbeel, and Jimmy Ba · 2019
Later among the works it cites.
Introduction to machine learning, Third edition
Ethem Alpaydin · 2020
Later among the works it cites.
The value equivalence principle for model-based reinforcement learning
Christopher Grimm, André Barreto, Satinder Singh, and David Silver · 2020
Later among the works it cites.
Mastering atari with discrete world models
Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi, and Jimmy Ba · 2020
Later among the works it cites.
Meta-learning in neural networks: A survey
Timothy Hospedales, Antreas Antoniou, Paul Micaelli, and Amos Storkey · 2020
Later among the works it cites.
Learning to Play: Reinforcement Learning and Games
Aske Plaat · 2020
Later among the works it cites.
From Chess and Atari to StarCraft and Beyond: How Game AI is Driving the World of AI
Sebastian Risi and Mike Preuss · 2020
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2020
Later among the works it cites.
Planning to explore via self-supervised world models
Ramanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel, Danijar Hafner, and Deepak Pathak · 2020
Later among the works it cites.
A survey of deep meta-learning
Mike Huisman, Jan N. van Rijn, and Aske Plaat · 2021
Closest in time.
Multiagent deep reinforcement learning: Challenges and directions towards human-like approaches
Annie Wong, Thomas Bäck, Anna V. Kononova, and Aske Plaat · 2021
Closest in time.