Fetching the paper…
Reading the bibliography…
Hierarchical Reinforcement Learning has been previously shown to speed up the convergence rate of RL planning algorithms as well as mitigate feature-based model misspecification (Mankowitz et.
Toward a theory of situation awareness in dynamic systems
Endsley, Mica R · 1995
Earlier work this paper cites.
Situation awareness is adaptive, externally directed consciousness
Smith, Kip and Hancock, Peter A · 1995
Earlier work this paper cites.
Lifelong robot learning
Thrun, Sebastian and Mitchell, Tom M · 1995
Earlier work this paper cites.
Stochastic approximation with two time scales
Borkar, Vivek S · 1997
Earlier work this paper cites.
Multi-time models for temporally abstract planning
Precup, Doina and Sutton, Richard S · 1997
Earlier work this paper cites.
Controlled markov chains with exponential risk-sensitive criteria: modularity, structured policies and applications
Avila-Godoy, Guadalupe and Fernández-Gaucherand, Emmanuel · 1998
Earlier work this paper cites.
Planning with macro-actions: Effect of initial value function estimate on convergence rate of value iteration
Hauskrecht, Milos · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, Richard and Barto, Andrew · 1998
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Sutton, Richard S, Precup, Doina, and Singh, Satinder · 1999
Earlier work this paper cites.
Policyblocks: An algorithm for creating useful macro-actions in reinforcement learning
Pickett, Marc and Barto, Andrew G · 2002
Cited alongside, same era.
A hierarchical type-2 fuzzy logic control architecture for autonomous mobile robots
Hagras, Hani et al · 2004
Cited alongside, same era.
Policy gradient methods for robotics
Peters, Jan and Schaal, Stefan · 2006
Cited alongside, same era.
Evaluation of the goal scoring patterns in european championship in portugal 2004
Yiannakos, A and Armatas, V · 2006
Cited alongside, same era.
Reinforcement learning of motor skills with policy gradients
Peters, Jan and Schaal, Stefan · 2008
Cited alongside, same era.
Probabilistic goal markov decision processes
Xu, Huan and Mannor, Shie · 2011
Pac-inspired option discovery in lifelong reinforcement learning
Brunskill, Emma and Li, Lihong · 2014
Later among the works it cites.
Time regularized interrupting options
Mankowitz, Daniel J, Mann, Timothy A, and Mannor, Shie · 2014
Later among the works it cites.
The option-critic architecture
Bacon, Pierre-Luc and Precup, Doina · 2015
Later among the works it cites.
Online planning for large markov decision processes with hierarchical decomposition
Bai, Aijun, Wu, Feng, and Chen, Xiaoping · 2015
Later among the works it cites.
Deep reinforcement learning in parameterized action space
Hausknecht, Matthew and Stone, Peter · 2015
Later among the works it cites.
Multirobot cooperative learning for semiautonomous control in urban search and rescue applications
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning parameterized skills
da Silva, B.C., Konidaris, G.D., and Barto, A.G · 2012
Cited alongside, same era.
The advantage of planning with options
Mann, Timothy A. and Mannor, Shie · 2013
Cited alongside, same era.
Helios base: An open source package for the robocup soccer 2d simulation
Akiyama, Hidehisa and Nakashima, Tomoharu · 2014
Cited alongside, same era.
Iterative Hierarchical Optimization for Misspecified Problems (IHOMP)
Mankowitz, Daniel J., Mann, Timothy A., and Mannor, Shie
Cited in the paper.
Adaptive Skills, Adaptive Partitions (ASAP)
Mankowitz, Daniel J., Mann, Timothy A., and Mannor, Shie
Cited in the paper.
Policy gradient for coherent risk measures
Tamar, Aviv, Chow, Yinlam, Ghavamzadeh, Mohammad, and Mannor, Shie
Cited in the paper.
Liu, Yugang and Nejat, Goldie · 2015
Later among the works it cites.
Learning when to switch between skills in a high dimensional domain
Mann, Timothy .A, Mankowitz Daniel J. Mannor Shie · 2015
Later among the works it cites.
Reinforcement learning with parameterized actions
Masson, Warwick and Konidaris, George · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A, Veness, Joel, Bellemare, Marc G, Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K, Ostrovski, Georg, et al · 2015
Later among the works it cites.