Fetching the paper…
Reading the bibliography…
An option is a short-term skill consisting of a control policy for a specified region of the state space, and a termination condition recognizing leaving that region.
Learning and executing generalized robot plans
R. E. Fikes, P. E. Hart, and N. J. Nilsson · 1972
Earlier work this paper cites.
L. E. Baum · 1972
Earlier work this paper cites.
A robust layered control system for a mobile robot
R. Brooks · 1986
Earlier work this paper cites.
Feudal reinforcement learning
P. Dayan and G. E. Hinton · 1992
Earlier work this paper cites.
Hierarchical control and learning for Markov decision processes
R. E. Parr · 1998
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
R. S. Sutton, D. Precup, and S. P. Singh · 1999
Earlier work this paper cites.
Learning attractor landscapes for learning motor primitives
A. Ijspeert, J. Nakanishi, and S. Schaal · 2002
Earlier work this paper cites.
Policy recognition in the abstract hidden Markov model
H. H. Bui, S. Venkatesh, and G. West · 2002
Earlier work this paper cites.
Discovering hierarchy in reinforcement learning with HEXQ
B. Hengst · 2002
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
A. G. Barto and S. Mahadevan · 2003
Earlier work this paper cites.
Optimization with EM and expectation-conjugate-gradient
R. Salakhutdinov, S. Roweis, and Z. Ghahramani · 2003
Earlier work this paper cites.
Hierarchical apprenticeship learning with application to quadruped locomotion
J. Z. Kolter, P. Abbeel, and A. Y. Ng · 2007
Cited alongside, same era.
Building portable options: Skill transfer in reinforcement learning
G. Konidaris and A. G. Barto · 2007
Cited alongside, same era.
The EM algorithm and extensions , volume 382
G. McLachlan and T. Krishnan · 2007
Cited alongside, same era.
Learning and generalization of motor skills by learning from demonstration
P. Pastor, H. Hoffmann, T. Asfour, and S. Schaal · 2009
Cited alongside, same era.
Hierarchical relative entropy policy search
C. Daniel, G. Neumann, and J. Peters · 2012
Cited alongside, same era.
Learning and generalization of complex tasks from unstructured demonstrations
S. Niekum, S. Osentoski, G. Konidaris, and A. Barto · 2012
Cited alongside, same era.
The option-critic architecture
P.-L. Bacon, J. Harb, and D. Precup · 2016
Later among the works it cites.
Probabilistic inference for determining options in reinforcement learning
C. Daniel, H. Van Hoof, J. Peters, and G. Neumann · 2016
Later among the works it cites.
Principled option learning in Markov decision processes
R. Fox, M. Moshkovitz, and N. Tishby · 2016
Later among the works it cites.
Hierarchical linearly-solvable Markov decision problems
A. Jonsson and V. Gómez · 2016
Later among the works it cites.
Multi-level discovery of deep options
R. Fox, S. Krishnan, I. Stoica, and K. Goldberg · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Robot learning from demonstration by constructing skill trees
G. Konidaris, S. Kuindersma, R. A. Grupen, and A. G. Barto · 2012
Cited alongside, same era.
Skills learning in robots by interaction with users and environment
S. Calinon · 2014
Cited alongside, same era.
Learning movement primitive attractor goals and sequential skills from kinesthetic demonstrations
S. Manschitz, J. Kober, M. Gienger, and J. Peters · 2015
Cited alongside, same era.
Transition state clustering: Unsupervised surgical trajectory segmentation for robot learning
S. Krishnan, A. Garg, S. Patil, C. Lea, G. Hager, P. Abbeel, and K. Goldberg · 2015
Cited alongside, same era.
Probabilistic segmentation applied to an assembly task
R. Lioutikov, G. Neumann, G. Maeda, and J. Peters · 2015
Cited alongside, same era.
Swirl: A sequential windowed inverse reinforcement learning algorithm for robot tasks with delayed rewards
S. Krishnan, A. Garg, R. Liaw, B. Thananjeyan, L. Miller, F. T. Pokorny, and K. Goldberg
Cited in the paper.
A. S. Vezhnevets, S. Osindero, T. Schaul, N. Heess, M. Jaderberg, D. Silver, and K. Kavukcuoglu · 2017
Closest in time.
Multi-modal imitation learning from unstructured demonstrations using generative adversarial nets
K. Hausman, Y. Chebotar, S. Schaal, G. Sukhatme, and J. Lim · 2017
Closest in time.
Robust imitation of diverse behaviors
Z. Wang, J. Merel, S. Reed, G. Wayne, N. de Freitas, and N. Heess · 2017
Closest in time.
Mutual alignment transfer learning
M. Wulfmeier, I. Posner, and P. Abbeel · 2017
Closest in time.
Cstochastic neural networks for hierarchical reinforcement learning
C. Florensa, Y. Duan, and P. Abbeel · 2017
Closest in time.