Fetching the paper…
Reading the bibliography…
This paper investigates a new approach to model-based reinforcement learning using background planning: mixing (approximate) dynamic programming updates and model-free updates, similar to the Dyna architecture.
A Generalized Inverse for Matrices
Roger Penrose · 1955
Earlier work this paper cites.
Integrated modeling and control based on reinforcement learning and dynamic programming
Richard S. Sutton · 1991
Earlier work this paper cites.
Feudal reinforcement learning
Peter Dayan and Geoffrey E Hinton · 1992
Earlier work this paper cites.
Scaling reinforcement learning algorithms by learning variable temporal resolution models
Satinder P Singh · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Prioritized Sweeping: Reinforcement Learning with Less Data and Less Time
Andrew W. Moore and Christopher G. Atkeson · 1993
Earlier work this paper cites.
Policy Invariance under Reward Transformations: Theory and Application to Reward Shaping
Andrew Y. Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Between MDPs and Semi-MDPs: A Framework for Temporal Abstraction in Reinforcement Learning
Richard S. Sutton, Doina Precup, and Satinder P. Singh · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
Thomas G Dietterich · 2000
Earlier work this paper cites.
Automatic Discovery of Subgoals in Reinforcement Learning using Diverse Density
Amy McGovern and Andrew G. Barto · 2001
Earlier work this paper cites.
Learning Options in Reinforcement Learning
Martin Stolle and Doina Precup · 2002
Earlier work this paper cites.
Intrinsically Motivated Reinforcement Learning
Satinder Singh, Andrew Barto, and Nuttapong Chentanez · 2004
Earlier work this paper cites.
Hierarchical dynamic programming for robot path planning
Bram Bakker, Zoran Zivkovic, and Ben Krose · 2005
Earlier work this paper cites.
Prioritization Methods for Accelerating MDP Solvers
David Wingate, Kevin D. Seppi, and Cs Byu Edu · 2005
Earlier work this paper cites.
A hierarchical approach to efficient reinforcement learning in deterministic domains
Carlos Diuk, Alexander L Strehl, and Michael L Littman · 2006
Earlier work this paper cites.
Sample-Based Learning and Search with Permanent and Transient Memories
David Silver, Richard S. Sutton, and Martin Müller · 2008
Earlier work this paper cites.
Skill Discovery in Continuous Reinforcement Learning Domains using Skill Chaining
George D. Konidaris and Andrew G. Barto · 2009
Earlier work this paper cites.
Learning Methods to Generate Good Plans: Integrating HTN Learning and Reinforcement Learning
Chad Hogg, U. Kuter, and Hector Muñoz-Avila · 2010
Earlier work this paper cites.
Double q-learning
Hado van Hasselt · 2010
Earlier work this paper cites.
Combined Task and Motion Planning for Mobile Manipulation
Jason Wolfe, Bhaskara Marthi, and Stuart Russell · 2010
Earlier work this paper cites.
Horde: A Scalable Real-Time Architecture for Learning Knowledge from Unsupervised Sensorimotor Interaction
Richard S. Sutton, Joseph Modayil, Michael Delp, Thomas Degris, Patrick M Pilarski, Adam White, and Doina Precup · 2011
Earlier work this paper cites.
Hierarchical solution of markov decision processes using macro-actions
Milos Hauskrecht, Nicolas Meuleau, Leslie Pack Kaelbling, Thomas L Dean, and Craig Boutilier · 2013
Earlier work this paper cites.
PAC-inspired Option Discovery in Lifelong Reinforcement Learning
Emma Brunskill and Lihong Li · 2014
Earlier work this paper cites.
Constructing symbolic representations for high-level planning
George Konidaris, Leslie Kaelbling, and Tomas Lozano-Perez · 2014
Earlier work this paper cites.
Scaling up approximate value iteration with options: Better policies with fewer iterations
Timothy Mann and Shie Mannor · 2014
Earlier work this paper cites.
Model Regularization for Stable Sample Roll-Outs
Erik Talvitie · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Approximate Value Iteration with Temporally Extended Actions
Timothy A. Mann, Shie Mannor, and Doina Precup · 2015
Cited alongside, same era.
Universal Value Function Approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Cited alongside, same era.
Improving Multi-Step Prediction of Learned Time Series Models
Arun Venkatraman, Martial Hebert, and J. Andrew Bagnell · 2015
Cited alongside, same era.
Constructing abstraction hierarchies using a skill-symbol loop
George Konidaris · 2016
Cited alongside, same era.
An Emphatic Approach to the Problem of Off-Policy Temporal-Difference Learning
Richard S. Sutton, Rupam A. Mahmood, and Martha White · 2016
Planning with Goal-Conditioned Policies
Soroush Nasiriany, Vitchyr Pong, Steven Lin, and Sergey Levine · 2019
Later among the works it cites.
Hill Climbing on Value Estimates for Search-Control in Dyna
Yangchen Pan, Hengshuai Yao, Amir-Massoud Farahmand, and Martha White · 2019
Later among the works it cites.
When to use Parametric Models in Reinforcement Learning?
Hado van Hasselt, Matteo Hessel, and John Aslanides · 2019
Later among the works it cites.
Planning with Expectation Models
Yi Wan, Muhammad Zaheer, Adam White, Martha White, and Richard S. Sutton · 2019
Later among the works it cites.
Model-Based Reinforcement Learning with Value-Targeted Regression
Alex Ayoub, Zeyu Jia, Csaba Szepesvári, Mengdi Wang, and Lin Yang · 2020
Later among the works it cites.
Fast reinforcement learning with generalized policy updates
André Barreto, Shaobo Hou, Diana Borsa, David Silver, and Doina Precup · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Value Iteration Networks
Aviv Tamar, Yi Wu, Garrett Thomas, Sergey Levine, and Pieter Abbeel · 2016
Cited alongside, same era.
Deep Reinforcement Learning with Double Q-learning
Hado van Hasselt, Arthur Guez, and David Silver · 2016
Cited alongside, same era.
Successor features for transfer in reinforcement learning
André Barreto, Will Dabney, Rémi Munos, Jonathan J. Hunt, Tom Schaul, David Silver, and Hado van Hasselt · 2017
Cited alongside, same era.
Value-Aware Loss Function for Model-based Reinforcement Learning
Amir-massoud Farahmand, Andre M S Barreto, and Daniel N Nikovski · 2017
Cited alongside, same era.
Planning with abstract markov decision processes
Nakul Gopalan, Michael Littman, James MacGlashan, Shawn Squire, Stefanie Tellex, John Winder, Lawson Wong, et al · 2017
Cited alongside, same era.
Value Prediction Network
Junhyuk Oh, Satinder Singh, and Honglak Lee · 2017
Cited alongside, same era.
Later among the works it cites.
Forethought and hindsight in credit assignment
Veronica Chelu, Doina Precup, and Hado P van Hasselt · 2020
Later among the works it cites.
Sparse Graphical Memory for Robust Planning
Scott Emmons, Ajay Jain, Misha Laskin, Thanard Kurutach, Pieter Abbeel, and Deepak Pathak · 2020
Later among the works it cites.
From Importance Sampling to Doubly Robust Policy Gradient
Jiawei Huang and Nan Jiang · 2020
Later among the works it cites.
Hallucinating Value: A Pitfall of Dyna-style Planning with Imperfect Environment Models
Taher Jafferjee, Ehsan Imani, Erin Talvitie, Martha White, and Micheal Bowling · 2020
Later among the works it cites.
What can I do here? A Theory of Affordances in Reinforcement Learning
Khimya Khetarpal, Zafarali Ahmed, Gheorghe Comanici, David Abel, and Doina Precup · 2020
Later among the works it cites.
Empirical Design in Reinforcement Learning
Andrew Patterson, Samuel Neumann, Martha White, and Adam White · 2020
Later among the works it cites.
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, Timothy Lillicrap, and David Silver · 2020
Later among the works it cites.
Planning in Hierarchical Reinforcement Learning: Guarantees for Using Local Policies
Tom Zahavy, Avinatan Hasidim, Haim Kaplan, and Yishay Mansour · 2020
Later among the works it cites.
Generating Adjacency-Constrained Subgoals in Hierarchical Reinforcement Learning
Tianren Zhang, Shangqi Guo, Tian Tan, Xiaolin Hu, and Feng Chen · 2020
Later among the works it cites.
DisTop: Discovering a Topological representation to learn diverse and rewarding skills
Arthur Aubret, Laetitia Matignon, and Salima Hassas · 2021
Later among the works it cites.
SNAP:Successor Entropy based Incremental Subgoal Discovery for Adaptive Navigation
Rohit K. Dubey, Samuel S. Sohn, Jimmy Abualdenien, Tyler Thrash, Christoph Hoelscher, André Borrmann, and Mubbasir Kapadia · 2021
Later among the works it cites.
Planning-Augmented Hierarchical Reinforcement Learning
Robert Gieselmann and Florian T. Pokorny · 2021
Later among the works it cites.
Successor Feature Landmarks for Long-Horizon Goal-Conditioned Reinforcement Learning
Christopher Hoang, Sungryull Sohn, Jongwook Choi, Wilka Carvalho, and Honglak Lee · 2021
Later among the works it cites.
Landmark-Guided Subgoal Generation in Hierarchical Reinforcement Learning
Junsu Kim, Younggyo Seo, and Jinwoo Shin · 2021
Later among the works it cites.
Average-Reward Learning and Planning with Options
Yi Wan, Abhishek Naik, and Richard S. Sutton · 2021
Later among the works it cites.
World Model as a Graph: Learning Latent Landmarks for Planning
Lunjun Zhang, Ge Yang, and Bradly C. Stadie · 2021
Later among the works it cites.
Learning Neuro-Symbolic Relational Transition Models for Bilevel Planning
Rohan Chitnis, Tom Silver, Joshua B. Tenenbaum, Tomás Lozano-Pérez, and Leslie Pack Kaelbling · 2022
Closest in time.
Investigating Compounding Prediction Errors in Learned Dynamics Models
Nathan Lambert, Kristofer Pister, and Roberto Calandra · 2022
Closest in time.
Reward-Respecting Subtasks for Model-Based Reinforcement Learning
Richard S. Sutton, Marlos C. Machado, G. Zacharias Holland, David Szepesvári, Finbarr Timbers, Brian Tanner, and Adam White · 2022
Closest in time.
Mastering Diverse Domains through World Models
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap · 2023
Closest in time.