Fetching the paper…
Reading the bibliography…
Model-based reinforcement learning is an appealing framework for creating agents that learn, plan, and act in sequential environments.
Correction to a formal basis for the heuristic determination of minimum cost paths
Peter E Hart, Nils J Nilsson, and Bertram Raphael · 1972
Earlier work this paper cites.
Science and statistics
George EP Box · 1976
Earlier work this paper cites.
Maximum likelihood from incomplete data via the em algorithm
Arthur P Dempster, Nan M Laird, and Donald B Rubin · 1977
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S Sutton · 1990
Earlier work this paper cites.
Prioritized sweeping: Reinforcement learning with less data and less time
Andrew W Moore and Christopher G Atkeson · 1993
Earlier work this paper cites.
On-line Q-learning using connectionist systems
GA Rummery and M Niranjan · 1994
Earlier work this paper cites.
TD models: Modeling the world at a mixture of time scales
Richard S Sutton · 1995
Earlier work this paper cites.
Reinforcement learning with replacing eligibility traces
Satinder P Singh and Richard S Sutton · 1996
Earlier work this paper cites.
Multi-time models for temporally abstract planning
Doina Precup and Richard S Sutton · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Rademacher and Gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael J. Kearns and Satinder P. Singh · 2002
Earlier work this paper cites.
The boosting approach to machine learning: An overview
Robert E Schapire · 2003
Earlier work this paper cites.
Ensemble selection from libraries of models
Rich Caruana, Alexandru Niculescu-Mizil, Geoff Crew, and Alex Ksikes · 2004
Earlier work this paper cites.
Using inaccurate models in reinforcement learning
Pieter Abbeel, Morgan Quigley, and Andrew Y Ng · 2006
Earlier work this paper cites.
Bandit based Monte-Carlo planning
Levente Kocsis and Csaba Szepesvári · 2006
Earlier work this paper cites.
Dyna-style planning with linear function approximation and prioritized sweeping
Richard S. Sutton, Csaba Szepesvári, Alborz Geramifard, and Michael H. Bowling · 2008
Earlier work this paper cites.
Model-based reinforcement learning with nearly tight exploration complexity bounds
István Szita and Csaba Szepesvári · 2010
Earlier work this paper cites.
PILCO: A model-based and data-efficient approach to policy search
Marc Peter Deisenroth and Carl Edward Rasmussen · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Cited alongside, same era.
On the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar, Rémi Munos, and Hilbert Kappen · 2012
Cited alongside, same era.
Foundations of Machine Learning
Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar · 2012
Cited alongside, same era.
Compositional planning using optimal option models
David Silver and Kamil Ciosek · 2012
Cited alongside, same era.
‘All models are wrong…’: An introduction to model uncertainty
Ernst Wit, Edwin van den Heuvel, and Jan-Willem Romeijn · 2012
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Artificial intelligence: a modern approach
Stuart J Russell and Peter Norvig · 2016
Later among the works it cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, et al · 2016
Later among the works it cites.
Cameron Allen, Kavosh Asadi, Melrose Roderick, Abdel-rahman Mohamed, George Konidaris, and Michael Littman · 2017
Later among the works it cites.
Multi-step reinforcement learning: A unifying algorithm
Kristopher De Asis, J. Fernando Hernandez-Garcia, G. Zacharias Holland, and Richard S. Sutton · 2017
Later among the works it cites.
Safe model-based reinforcement learning with stability guarantees
Felix Berkenkamp, Matteo Turchetta, Angela Schoellig, and Andreas Krause · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Probability in Banach Spaces: isoperimetry and processes
Michel Ledoux and Michel Talagrand · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Cited alongside, same era.
Learning neural network policies with guided policy search under unknown dynamics
Sergey Levine and Pieter Abbeel · 2014
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Cited alongside, same era.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Cited alongside, same era.
Model regularization for stable sample rollouts
Erik Talvitie · 2014
Cited alongside, same era.
Thomas M Moerland, Joost Broekens, and Catholijn M Jonker · 2017
Later among the works it cites.
The predictron: End-to-end learning and planning
David Silver, Hado van Hasselt, Matteo Hessel, Tom Schaul, Arthur Guez, Tim Harley, Gabriel Dulac-Arnold, David Reichert, Neil Rabinowitz, Andre Barreto, et al · 2017
Later among the works it cites.
Self-correcting models for model-based reinforcement learning
Erik Talvitie · 2017
Later among the works it cites.
Equivalence between Wasserstein and value-aware model-based reinforcement learning
Kavosh Asadi, Evan Cater, Dipendra Misra, and Michael L. Littman · 2018
Later among the works it cites.
Lipschitz continuity in model-based reinforcement learning
Kavosh Asadi, Dipendra Misra, and Michael L. Littman · 2018
Later among the works it cites.
Sample-efficient deep RL with generative adversarial tree search
Kamyar Azizzadenesheli, Brandon Yang, Weitang Liu, Emma Brunskill, Zachary C. Lipton, and Animashree Anandkumar · 2018
Later among the works it cites.
Improving pilco with bayesian neural network dynamics models
Yarin Gal, Rowan McAllister, and Carl Edward Rasmussen · 2018
Later among the works it cites.
Model-ensemble trust-region policy optimization
Thanard Kurutach, Ignasi Clavera, Yan Duan, Aviv Tamar, and Pieter Abbeel · 2018
Later among the works it cites.
On value function representation of long horizon problems
Lucas Lehnert, Romain Laroche, and Harm van Seijen · 2018
Later among the works it cites.
Algorithmic framework for model-based deep reinforcement learning with theoretical guarantees
Yuping Luo, Huazhe Xu, Yuanzhi Li, Yuandong Tian, Trevor Darrell, and Tengyu Ma · 2018
Later among the works it cites.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Marlos C Machado, Marc G Bellemare, Erik Talvitie, Joel Veness, Matthew Hausknecht, and Michael Bowling · 2018
Later among the works it cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Anusha Nagabandi, Gregory Kahn, Ronald S Fearing, and Sergey Levine · 2018
Later among the works it cites.
Model-based active exploration
Pranav Shyam, Wojciech Jaskowski, and Faustino Gomez · 2018
Later among the works it cites.
Model-based reinforcement learning in contextual decision processes
Wen Sun, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
DeepMDP: Learning continuous latent space models for representation learning
Carles Gelada, Saurabh Kumar, Jacob Buckman, Ofir Nachum, and Marc G. Bellemare · 2019
Closest in time.