Fetching the paper…
Reading the bibliography…
Learning models of the environment from data is often viewed as an essential component to building intelligent reinforcement learning (RL) agents.
Learning Causal State Representations of Partially Observable Environments
Amy Zhang, Zachary C Lipton, Luis Pineda, Kamyar Azizzadenesheli, Anima Anandkumar, Laurent Itti, Joelle Pineau, and Tommaso Furlanello · 1906
Earlier work this paper cites.
Neuronlike Adaptive Elements That Can Solve Difficult Learning Control Problems
Andrew G. Barto, Richard S. Sutton, and Charles W Anderson · 1983
Earlier work this paper cites.
Learning to Predict by the Methods of Temporal Differences
Richard S. Sutton · 1988
Earlier work this paper cites.
Markov Decision Processes—Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Neuro-Dynamic Programming
Dimitri P. Bertsekas and John N. Tsitsiklis · 1996
Earlier work this paper cites.
Model Minimization in Markov Decision Processes
Thomas Dean and Robert Givan · 1997
Earlier work this paper cites.
Between MDPs and Semi-MDPs: A Framework for Temporal Abstraction in Reinforcement Learning
Richard S Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Value-Directed Compression of POMDPs
Pascal Poupart and Craig Boutilier · 2002
Earlier work this paper cites.
Equivalence Notions and Model Minimization in Markov Decision Processes
Robert Givan, Thomas Dean, and Matthew Greig · 2003
Earlier work this paper cites.
Artificial Intelligence: A Modern Approach
Stuart J. Russell and Peter Norvig · 2003
Earlier work this paper cites.
Metrics for Finite Markov Decision Processes
Norm Ferns, Prakash Panangaden, and Doina Precup · 2004
Earlier work this paper cites.
Approximate Homomorphisms: A Framework for Non-Exact Minimization in Markov Decision Processes
Balaraman Ravindran and Andrew G Barto · 2004
Earlier work this paper cites.
Towards a Unified Theory of State Abstraction for MDPs
Lihong Li, Thomas J. Walsh, and Michael L. Littman · 2006
Earlier work this paper cites.
An Analysis of Linear Models, Linear Value-Function Approximation, and Feature Selection for Reinforcement Learning
Ronald Parr, Lihong Li, Gavin Taylor, Christopher Painter-Wakefield, and Michael L. Littman · 2008
Earlier work this paper cites.
Dyna-Style Planning with Linear Function Approximation and Prioritized Sweeping
Richard S. Sutton, Csaba Szepesvári, Alborz Geramifard, and Michael Bowling · 2008
Earlier work this paper cites.
Bounding Performance Loss in Approximate MDP Homomorphisms
Jonathan Taylor, Doina Precup, and Prakash Panagaden · 2009
Cited alongside, same era.
Algorithms for Reinforcement Learning
Csaba Szepesvári · 2010
Cited alongside, same era.
Maximum Likelihood Estimation and Inference
Russell B. Millar · 2011
Cited alongside, same era.
Value-Aware Loss Function for Model Learning in Reinforcement Learning
Amir-Massoud Farahmand, André Barreto, and Daniel Nikovski · 2013
Cited alongside, same era.
Reinforcement Learning with Misspecified Model Classes
Joshua Joseph, Alborz Geramifard, John W Roberts, Jonathan P How, and Nicholas Roy · 2013
Cited alongside, same era.
Value-Directed Belief State Approximation for POMDPs
Pascal Poupart and Craig Boutilier · 2013
TreeQN and ATreeC: Differentiable Tree-Structured Models for Deep Reinforcement Learning
G Farquhar, T Rocktäschel, M Igl, and S Whiteson · 2018
Later among the works it cites.
Deep Variational Reinforcement Learning for POMDPs
Maximillian Igl, Luisa Zintgraf, Tuan Anh Le, Frank Wood, and Shimon Whiteson · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Later among the works it cites.
A Geometric Perspective on Optimal Representations for Reinforcement Learning
Marc Bellemare, Will Dabney, Robert Dadashi, Adrien Ali Taiga, Pablo Samuel Castro, Nicolas Le Roux, Dale Schuurmans, Tor Lattimore, and Clare Lyle · 2019
Later among the works it cites.
The Value Function Polytope in Reinforcement Learning
Robert Dadashi, Adrien Ali Taiga, Nicolas Le Roux, Dale Schuurmans, and Marc G. Bellemare · 2019
Later among the works it cites.
Combined Reinforcement Learning via Abstract Representations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Recurrent Models of Visual Attention
Volodymyr Mnih, Nicolas Heess, Alex Graves, and Koray Kavukcuoglu · 2014
Cited alongside, same era.
Embed to Control: A Locally Linear Latent Dynamics Model for Control from Raw Images
Manuel Watter, Jost Springenberg, Joschka Boedecker, and Martin Riedmiller · 2015
Cited alongside, same era.
Value Iteration Networks
Aviv Tamar, Yi Wu, Garrett Thomas, Sergey Levine, and Pieter Abbeel · 2016
Cited alongside, same era.
Value-Aware Loss Function for Model-Based Reinforcement Learning
Amir-Massoud Farahmand, André Barreto, and Daniel Nikovski · 2017
Cited alongside, same era.
Value Prediction Networks
Junhyuk Oh, Satinder Singh, and Honglak Lee · 2017
Cited alongside, same era.
The Predictron: End-to-End Learning and Planning
David Silver, Hado van Hasselt, Matteo Hessel, Tom Schaul, Arthur Guez, Tim Harley, Gabriel Dulac-Arnold, David Reichert, Neil Rabinowitz, Andre Barreto, et al · 2017
Cited alongside, same era.
Vincent François-Lavet, Yoshua Bengio, Doina Precup, and Joelle Pineau · 2019
Later among the works it cites.
DeepMDP: Learning Continuous Latent Space Models for Representation Learning
Carles Gelada, Saurabh Kumar, Jacob Buckman, Ofir Nachum, and Marc G Bellemare · 2019
Later among the works it cites.
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2019
Later among the works it cites.
Policy-Aware Model Learning for Policy Gradient Methods
Romina Abachi, Mohammad Ghavamzadeh, and Amir massoud Farahmand · 2020
Closest in time.
Model-Based Reinforcement Learning with Value-Targeted Regression
Alex Ayoub, Zeyu Jia, Csaba Szepesvári, Mengdi Wang, and Lin Yang · 2020
Closest in time.
Learning Discrete State Abstractions with Deep Variational Inference
Ondrej Biza, Robert Platt, Jan-Willem van de Meent, and Lawson LS Wong · 2020
Closest in time.
Scalable Methods for Computing State Similarity in Deterministic Markov Decision Processes
Pablo Samuel Castro · 2020
Closest in time.
The Value-Improvement Path: Towards Better Representations for Reinforcement Learning, 2020
Will Dabney, André Barreto, Mark Rowland, Robert Dadashi, John Quan, Marc G. Bellemare, and David Silver · 2020
Closest in time.
Plannable Approximations to MDP Homomorphisms: Equivariance under Actions
Elise van der Pol, Thomas Kipf, Frans A Oliehoek, and Max Welling · 2020
Closest in time.