Fetching the paper…
Reading the bibliography…
Hierarchical agents have the potential to solve sequential decision making tasks with greater sample efficiency than their non-hierarchical counterparts because hierarchical agents can break down tasks into sets of subtasks that only require short sequences of decisions.
Learning to generate sub-goals for action sequences
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
Q-learning
Christopher J. C. H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
HQ-learning
Marco Wiering and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Between MDPs and semi-MDPs: a framework for temporal abstraction in reinforcement learning
R. S. Sutton, D. Precup, and S. Singh · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the MAXQ value function decomposition
T.G. Dietterich · 2000
Earlier work this paper cites.
Automatic discovery of subgoals in reinforcement learning using diverse density
A. McGovern and A.G. Barto · 2001
Earlier work this paper cites.
Q-cut—dynamic discovery of sub-goals in reinforcement learning
I. Menache, S. Mannor, and N. Shimkin · 2002
Cited alongside, same era.
Hierarchical reinforcement learning with subpolicies specializing for learned subgoals
Bram Bakker and Jürgen Schmidhuber · 2004
Cited alongside, same era.
Identifying useful subgoals in reinforcement learning by local graph partitioning
Ö. Şimşek, A.P. Wolfe, and A.G. Barto · 2005
Cited alongside, same era.
Skill discovery in continuous reinforcement learning domains using skill chaining
G.D. Konidaris and A.G. Barto · 2009
Cited alongside, same era.
MuJoCo: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Cited alongside, same era.
Universal value function approximators
T. Schaul, D. Horgan, K. Gregor, and D. Silver · 2015
Later among the works it cites.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
T.D. Kulkarni, K. Narasimhan, A. Saeedi, and J. Tenenbaum · 2016
Later among the works it cites.
Hindsight experience replay
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, P. Abbeel, and W. Zaremba · 2017
Closest in time.
The option-critic architecture
P-L Bacon, J. Harb, and D. Precup · 2017
Closest in time.
FeUdal networks for hierarchical reinforcement learning
A. Vezhnevets, S. Osindero, T. Schaul, N. Heess, M. Jaderberg, D. Silver, and K. Kavukcuoglu · 2017
Closest in time.
Data-efficient hierarchical reinforcement learning
O. Nachum, S. Gu, H. Lee, and S. Levine · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T.P. Lillicrap, J.J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Cited alongside, same era.