Fetching the paper…
Reading the bibliography…
This paper presents a way of solving Markov Decision Processes that combines state abstraction and temporal abstraction.
Dynamic Programming
R. Bellman · 1957
Earlier work this paper cites.
Learning to Solve Problems by Searching for Macro-Operators
R. Korf · 1985
Earlier work this paper cites.
Complete solution of the eight-puzzle and the benefit of node ordering in ida*
A. Reinefeld · 1993
Earlier work this paper cites.
TD Models: Modeling the World at a Mixture of Time Scales
R. S. Sutton · 1995
Earlier work this paper cites.
The MAXQ Method for Hierarchical Reinforcement Learning
T. G. Dietterich · 1998
Earlier work this paper cites.
Theoretical results on reinforcement learning with temporally abstract options
D. Precup, R. S. Sutton, and S. Singh · 1998
Earlier work this paper cites.
Between MDPs and Semi-MDPs: A Framework for Temporal Abstraction in Reinforcement Learning
R. Sutton, D. Precup, and S. Singh · 1999
Earlier work this paper cites.
On the existence of fixed points for approximate value iteration and temporal-difference learning
D. P. De Farias and B. Van Roy · 2000
Cited alongside, same era.
Automated state abstraction for options using the U-tree algorithm
A. Jonsson and A. G. Barto · 2001
Cited alongside, same era.
State abstraction for programmable reinforcement learning agents
D. Andre and S. J. Russell · 2002
Cited alongside, same era.
Discovering hierarchy in reinforcement learning with HEXQ
B. Hengst · 2002
Cited alongside, same era.
State Abstraction Discovery from Irrelevant State Variables.
N. K. Jong and P. Stone · 2005
Cited alongside, same era.
TD(0) Leads to Better Policies than Approximate Value Iteration
B. Van Roy · 2005
Cited alongside, same era.
An analysis of linear models, linear value-function approximation, and feature selection for reinforcement learning
R. Parr, L. Li, G. Taylor, C. Painter-Wakefield, and M. L. Littman · 2008
Later among the works it cites.
Linear options
J. Sorg and S. Singh · 2010
Later among the works it cites.
Convergent fitted value iteration with linear function approximation
D. J. Lizotte · 2011
Later among the works it cites.
A neural signature of hierarchical reinforcement learning
J. J. Ribas-Fernandes, A. Solway, C. Diuk, J. T. McGuire, A. G. Barto, Y. Niv, and M. M. Botvinick · 2011
Later among the works it cites.
Dynamic Programming and Optimal Control
D. P. Bertsekas · 2012
Later among the works it cites.
Compositional planning using optimal option models
D. Silver and K. Ciosek · 2012
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Building Portable Options: Skill Transfer in Reinforcement Learning.
G. Konidaris and A. G. Barto · 2007
Cited alongside, same era.
Notes on the “15” puzzle
W. E. Story
Cited in the paper.
Automatic discovery and transfer of maxq hierarchies in a complex system
H. Wang, W. Li, and X. Zhou · 2012
Later among the works it cites.