Fetching the paper…
Reading the bibliography…
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, Ronald J · 1992
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
Parr, Ron and Russell, Stuart · 1998
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Sutton, Richard S, Precup, Doina, and Singh, Satinder · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
Dietterich, Thomas G · 2000
Earlier work this paper cites.
Temporal abstraction in reinforcement learning
Precup, Doina · 2000
Earlier work this paper cites.
Programmable reinforcement learning agents
Andre, David and Russell, Stuart · 2001
Earlier work this paper cites.
State abstraction for programmable reinforcement learning agents
Andre, David and Russell, Stuart · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Kearns, Michael and Singh, Satinder · 2002
Earlier work this paper cites.
Q-cut—dynamic discovery of sub-goals in reinforcement learning
Menache, Ishai, Mannor, Shie, and Shimkin, Nahum · 2002
Earlier work this paper cites.
Learning options in reinforcement learning
Stolle, Martin and Precup, Doina · 2002
Earlier work this paper cites.
Hierarchical reinforcement learning based on subgoal discovery and subpolicy specialization
Bakker, Bram and Schmidhuber, Jürgen · 2004
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Greensmith, Evan, Bartlett, Peter L, and Baxter, Jonathan · 2004
Earlier work this paper cites.
Concurrent hierarchical reinforcement learning
Marthi, Bhaskara, Lantham, David, Guestrin, Carlos, and Russell, Stuart · 2004
Cited alongside, same era.
Building portable options: Skill transfer in reinforcement learning
Konidaris, George and Barto, Andrew G · 2007
Cited alongside, same era.
Using motion primitives in probabilistic sample-based planning for humanoid robots
Hauser, Kris, Bretl, Timothy, Harada, Kensuke, and Latombe, Jean-Claude · 2008
Cited alongside, same era.
Curriculum learning
Bengio, Yoshua, Louradour, Jérôme, Collobert, Ronan, and Weston, Jason · 2009
Cited alongside, same era.
Reinforcement learning for mapping instructions to actions
Branavan, S.R.K., Chen, Harr, Zettlemoyer, Luke S., and Barzilay, Regina · 2009
Cited alongside, same era.
Learning to follow navigational directions
Vogel, Adam and Jurafsky, Dan · 2010
Cited alongside, same era.
RMSProp (unpublished), 2012
Tieleman, Tijmen · 2012
Later among the works it cites.
Weakly supervised learning of semantic parsers for mapping instructions to actions
Artzi, Yoav and Zettlemoyer, Luke · 2013
Later among the works it cites.
A neural network for factoid question answering over paragraphs
Iyyer, Mohit, Boyd-Graber, Jordan, Claudino, Leonardo, Socher, Richard, and Daumé III, Hal · 2014
Later among the works it cites.
The option-critic architecture
Bacon, Pierre-Luc and Precup, Doina · 2015
Later among the works it cites.
Neural programmer: Inducing latent programs with gradient descent
Neelakantan, Arvind, Le, Quoc V, and Sutskever, Ilya · 2015
Later among the works it cites.
Learning grounded finite-state representations from unstructured demonstrations
Niekum, Scott, Osentoski, Sarah, Konidaris, George, Chitta, Sachin, Marthi, Bhaskara, and Barto, Andrew G · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning to interpret natural language navigation instructions from observations
Chen, David L. and Mooney, Raymond J · 2011
Cited alongside, same era.
Robot learning from demonstration by constructing skill trees
Konidaris, George, Kuindersma, Scott, Grupen, Roderic, and Barto, Andrew · 2011
Cited alongside, same era.
Understanding natural language commands for robotic navigation and mobile manipulation
Tellex, Stefanie, Kollar, Thomas, Dickerson, Steven, Walter, Matthew R., Banerjee, Ashis Gopal, Teller, Seth, and Roy, Nicholas · 2011
Cited alongside, same era.
Hierarchical relative entropy policy search
Daniel, Christian, Neumann, Gerhard, and Peters, Jan · 2012
Cited alongside, same era.
Semantic compositionality through recursive matrix-vector spaces
Socher, Richard, Huval, Brody, Manning, Christopher, and Ng, Andrew · 2012
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
Schulman, John, Moritz, Philipp, Levine, Sergey, Jordan, Michael, and Abbeel, Pieter
Cited in the paper.
Later among the works it cites.
Learning to compose neural networks for question answering
Andreas, Jacob, Rohrbach, Marcus, Darrell, Trevor, and Klein, Dan · 2016
Closest in time.
Learning modular neural network policies for multi-task and multi-robot transfer
Devin, Coline, Gupta, Abhishek, Darrell, Trevor, Abbeel, Pieter, and Levine, Sergey · 2016
Closest in time.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Kulkarni, Tejas D, Narasimhan, Karthik R, Saeedi, Ardavan, and Tenenbaum, Joshua B · 2016
Closest in time.
Neural programmer-interpreters
Reed, Scott and de Freitas, Nando · 2016
Closest in time.
Strategic attentive writer for learning macro-actions
Vezhnevets, Alexander, Mnih, Volodymyr, Agapiou, John, Osindero, Simon, Graves, Alex, Vinyals, Oriol, and Kavukcuoglu, Koray · 2016
Closest in time.