Fetching the paper…
Reading the bibliography…
One of the main challenges in model-based reinforcement learning (RL) is to decide which aspects of the environment should be modeled.
Learning causal state representations of partially observable environments
Amy Zhang, Zachary C Lipton, Luis Pineda, Kamyar Azizzadenesheli, Anima Anandkumar, Laurent Itti, Joelle Pineau, and Tommaso Furlanello · 1906
Earlier work this paper cites.
Markov Decision Processes—Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
TD models: Modeling the world at a mixture of time scales
Richard S. Sutton · 1995
Earlier work this paper cites.
Model minimization in markov decision processes
Thomas Dean and Robert Givan · 1997
Earlier work this paper cites.
Equivalence notions and model minimization in markov decision processes
Robert Givan, Thomas Dean, and Matthew Greig · 2003
Earlier work this paper cites.
Value-directed compression of pomdps
Pascal Poupart, Craig Boutilier, et al · 2003
Earlier work this paper cites.
Artificial Intelligence: A Modern Approach
Stuart J. Russell and Peter Norvig · 2003
Earlier work this paper cites.
Metrics for finite markov decision processes
Norm Ferns, Prakash Panangaden, and Doina Precup · 2004
Earlier work this paper cites.
Approximate homomorphisms: A framework for non-exact minimization in markov decision processes
Balaraman Ravindran and Andrew G Barto · 2004
Earlier work this paper cites.
Efficient selectivity and backup operators in monte-carlo tree search
Rémi Coulom · 2006
Earlier work this paper cites.
Bounding performance loss in approximate mdp homomorphisms
Jonathan Taylor, Doina Precup, and Prakash Panagaden · 2008
Earlier work this paper cites.
Toward a unified theory of development
John P Spencer, Michael SC Thomas, and JL McClelland · 2009
Earlier work this paper cites.
Algorithms for Reinforcement Learning
Csaba Szepesvári · 2010
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Reinforcement learning with misspecified model classes
Joshua Joseph, Alborz Geramifard, John W Roberts, Jonathan P How, and Nicholas Roy · 2013
Cited alongside, same era.
Value-directed belief state approximation for POMDPs
Pascal Poupart and Craig Boutilier · 2013
Cited alongside, same era.
Embed to control: a locally linear latent dynamics model for control from raw images
Manuel Watter, Jost Tobias Springenberg, Joschka Boedecker, and Martin Riedmiller · 2015
Cited alongside, same era.
Value iteration networks
Aviv Tamar, Yi Wu, Garrett Thomas, Sergey Levine, and Pieter Abbeel · 2016
Cited alongside, same era.
Value-Aware Loss Function for Model-based Reinforcement Learning
Amir-Massoud Farahmand, André Barreto, and Daniel Nikovski · 2017
Cited alongside, same era.
Value prediction network
Junhyuk Oh, Satinder Singh, and Honglak Lee · 2017
Combined reinforcement learning via abstract representations
Vincent François-Lavet, Yoshua Bengio, Doina Precup, and Joelle Pineau · 2019
Later among the works it cites.
Deepmdp: Learning continuous latent space models for representation learning
Carles Gelada, Saurabh Kumar, Jacob Buckman, Ofir Nachum, and Marc G Bellemare · 2019
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2019
Later among the works it cites.
Model-based reinforcement learning with value-targeted regression
Alex Ayoub, Zeyu Jia, Csaba Szepesvari, Mengdi Wang, and Lin Yang · 2020
Later among the works it cites.
Learning discrete state abstractions with deep variational inference
Ondrej Biza, Robert Platt, Jan-Willem van de Meent, and Lawson LS Wong · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The predictron: End-to-end learning and planning
David Silver, Hado van Hasselt, Matteo Hessel, Tom Schaul, Arthur Guez, Tim Harley, Gabriel Dulac-Arnold, David Reichert, Neil Rabinowitz, Andre Barreto, et al · 2017
Cited alongside, same era.
Efficient model-based deep reinforcement learning with variational state tabulation
Dane Corneil, Wulfram Gerstner, and Johanni Brea · 2018
Cited alongside, same era.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Vlad Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al · 2018
Cited alongside, same era.
Iterative value-aware model learning
Amir-massoud Farahmand · 2018
Cited alongside, same era.
Treeqn and atreec: Differentiable tree-structured models for deep reinforcement learning
G Farquhar, T Rocktäschel, M Igl, and S Whiteson · 2018
Cited alongside, same era.
Deep variational reinforcement learning for pomdps
Maximilian Igl, Luisa Zintgraf, Tuan Anh Le, Frank Wood, and Shimon Whiteson · 2018
Cited alongside, same era.
Later among the works it cites.
Scalable methods for computing state similarity in deterministic markov decision processes
Pablo Samuel Castro · 2020
Later among the works it cites.
The value equivalence principle for model-based reinforcement learning
Christopher Grimm, Andre Barreto, Satinder Singh, and David Silver · 2020
Later among the works it cites.
Sample complexity of reinforcement learning using linearly combined model ensembles
Aditya Modi, Nan Jiang, Ambuj Tewari, and Satinder Singh · 2020
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2020
Later among the works it cites.
Plannable approximations to mdp homomorphisms: Equivariance under actions
Elise van der Pol, Thomas Kipf, Frans A Oliehoek, and Max Welling · 2020
Later among the works it cites.
Learning invariant representations for reinforcement learning without reconstruction
Amy Zhang, Rowan McAllister, Roberto Calandra, Yarin Gal, and Sergey Levine · 2020
Later among the works it cites.
Contrastive behavioral similarity embeddings for generalization in reinforcement learning
Rishabh Agarwal, Marlos C Machado, Pablo Samuel Castro, and Marc G Bellemare · 2021
Closest in time.
Online and offline reinforcement learning by planning with a learned model
Julian Schrittwieser, Thomas Hubert, Amol Mandhane, Mohammadamin Barekatain, Ioannis Antonoglou, and David Silver · 2021
Closest in time.