Fetching the paper…
Reading the bibliography…
The composition of elementary behaviors to solve challenging transfer learning problems is one of the key elements in building intelligent machines.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Temporal abstraction in reinforcement learning
Doina Precup · 2000
Earlier work this paper cites.
Neurophysiological mechanisms underlying the understanding and imitation of action
Giacomo Rizzolatti, Leonardo Fogassi, and Vittorio Gallese · 2001
Earlier work this paper cites.
Learning movement primitives
Stefan Schaal, Jan Peters, Jun Nakanishi, and Auke Ijspeert · 2005
Earlier work this paper cites.
Towards cognitive robots: Building hierarchical task representations of manipulations from human demonstration
R Zoliner, Michael Pardowitz, Steffen Knoop, and Rüdiger Dillmann · 2005
Earlier work this paper cites.
Teaching sequential tasks with repetition through demonstration
Harini Veeraraghavan and Manuela Veloso · 2008
Earlier work this paper cites.
Compositionality of optimal control laws
Emanuel Todorov · 2009
Earlier work this paper cites.
Learning table tennis with a mixture of motor primitives
Katharina Muelling, Jens Kober, and Jan Peters · 2010
Earlier work this paper cites.
Learning parametric dynamic movement primitives from multiple demonstrations
Takamitsu Matsubara, Sang-Ho Hyon, and Jun Morimoto · 2011
Earlier work this paper cites.
Imitating others by composition of primitive actions: A neuro-dynamic model
Hiroaki Arie, Takafumi Arakaki, Shigeki Sugano, and Jun Tani · 2012
Earlier work this paper cites.
Robot learning from demonstration by constructing skill trees
George Konidaris, Scott Kuindersma, Roderic Grupen, and Andrew Barto · 2012
Cited alongside, same era.
Dynamical movement primitives: learning attractor models for motor behaviors
Auke Jan Ijspeert, Jun Nakanishi, Heiko Hoffmann, Peter Pastor, and Stefan Schaal · 2013
Cited alongside, same era.
Probabilistic movement primitives
Alexandros Paraschos, Christian Daniel, Jan R Peters, and Gerhard Neumann · 2013
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Building machines that learn and think like people
Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gershman · 2017
Later among the works it cites.
Show, attend and interact: Perceivable human-robot social interaction through neural attention q-network
Ahmed Hussain Qureshi, Yutaka Nakamura, Yuichiro Yoshikawa, and Hiroshi Ishiguro · 2017
Later among the works it cites.
Himanshu Sahni, Saurabh Kumar, Farhan Tejani, and Charles Isbell · 2017
Later among the works it cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Feudal networks for hierarchical reinforcement learning
Alexander Sasha Vezhnevets, Simon Osindero, Tom Schaul, Nicolas Heess, Max Jaderberg, David Silver, and Koray Kavukcuoglu · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Strategic attentive writer for learning macro-actions
Alexander Vezhnevets, Volodymyr Mnih, Simon Osindero, Alex Graves, Oriol Vinyals, John Agapiou, et al · 2016
Cited alongside, same era.
The option-critic architecture
Pierre-Luc Bacon, Jean Harb, and Doina Precup · 2017
Cited alongside, same era.
Stochastic neural networks for hierarchical reinforcement learning
Carlos Florensa, Yan Duan, and Pieter Abbeel · 2017
Cited alongside, same era.
Learning composable models of parameterized skills
Leslie Pack Kaelbling and Tomás Lozano-Pérez · 2017
Cited alongside, same era.
Composable deep reinforcement learning for robotic manipulation
Tuomas Haarnoja, Vitchyr Pong, Aurick Zhou, Murtaza Dalal, Pieter Abbeel, and Sergey Levine
Cited in the paper.
Later among the works it cites.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke van Hoof, and David Meger · 2018
Later among the works it cites.
When waiting is not an option: Learning options with a deliberation cost
Jean Harb, Pierre-Luc Bacon, Martin Klissarov, and Doina Precup · 2018
Later among the works it cites.
Data-efficient hierarchical reinforcement learning
Ofir Nachum, Shane Gu, Honglak Lee, and Sergey Levine · 2018
Later among the works it cites.
Intrinsically motivated reinforcement learning for human–robot interaction in the real-world
Ahmed Hussain Qureshi, Yutaka Nakamura, Yuichiro Yoshikawa, and Hiroshi Ishiguro · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Mcp: Learning composable hierarchical control with multiplicative compositional policies
Xue Bin Peng, Michael Chang, Grace Zhang, Pieter Abbeel, and Sergey Levine · 2019
Closest in time.