Fetching the paper…
Reading the bibliography…
We introduce FeUdal Networks (FuNs): a novel architecture for hierarchical reinforcement learning.
Spatial localization does not require the presence of local cues
Morris, Richard GM · 1981
Earlier work this paper cites.
A focused back-propagation algorithm for temporal pattern recognition
Mozer, Michael C · 1989
Earlier work this paper cites.
Neural sequence chunkers
Schmidhuber, Jürgen · 1991
Earlier work this paper cites.
Improving generalization for temporal difference learning: The successor representation
Dayan, Peter · 1993
Earlier work this paper cites.
Feudal reinforcement learning
Dayan, Peter and Hinton, Geoffrey E · 1993
Earlier work this paper cites.
Td models: Modeling the world at a mixture of time scales
Sutton, Richard S · 1995
Earlier work this paper cites.
Prioritized goal decomposition of markov decision processes: Toward a synthesis of classical and decision theoretic planning
Boutilier, Craig, Brafman, Ronen I, and Geib, Christopher · 1997
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Planning with closed-loop macro actions
Precup, Doina, Sutton, Richard S, and Singh, Satinder P · 1997
Earlier work this paper cites.
Hq-learning
Wiering, Marco and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
Parr, Ronald and Russell, Stuart · 1998
Earlier work this paper cites.
Theoretical results on reinforcement learning with temporally abstract options
Precup, Doina, Sutton, Richard S, and Singh, Satinder · 1998
Cited alongside, same era.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, Richard S, Precup, Doina, and Singh, Satinder · 1999
Cited alongside, same era.
Hierarchical reinforcement learning with the maxq value function decomposition
Dietterich, Thomas G · 2000
Cited alongside, same era.
Temporal abstraction in reinforcement learning
Precup, Doina · 2000
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
Bellemare, Marc G, Naddaf, Yavar, Veness, Joel, and Bowling, Michael · 2012
Cited alongside, same era.
Hierarchical learning in stochastic domains: Preliminary results
Kaelbling, Leslie Pack · 2014
Trust region policy optimization
Schulman, John, Levine, Sergey, Moritz, Philipp, Jordan, Michael I, and Abbeel, Pieter · 2015
Later among the works it cites.
Beattie, Charles, Leibo, Joel Z., Teplyashin, Denis, Ward, Tom, Wainwright, Marcus, Küttler, Heinrich, Lefrancq, Andrew, Green, Simon, Valdés, Víctor, Sadik, Amir, Schrittwieser, Julian, Anderson, Keith, York, Sarah, Cant, Max, Cain, Adam, Bolton, Adrian, Gaffney, Stephen, King, Helen, Hassabis, Demis, Legg, Shane, and Petersen, Stig · 2016
Later among the works it cites.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, Max, Mnih, Volodymyr, Czarnecki, Wojciech Marian, Schaul, Tom, Leibo, Joel Z, Silver, David, and Kavukcuoglu, Koray · 2016
Later among the works it cites.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Kulkarni, Tejas D., Narasimhan, Karthik R., Saeedi, Ardavan, and Tenenbaum, Joshua B · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A clockwork rnn
Koutník, Jan, Greff, Klaus, Gomez, Faustino, and Schmidhuber, Jürgen · 2014
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Levine, Sergey, Finn, Chelsea, Darrell, Trevor, and Abbeel, Pieter · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, Timothy P, Hunt, Jonathan J, Pritzel, Alexander, Heess, Nicolas, Erez, Tom, Tassa, Yuval, Silver, David, and Wierstra, Daan · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A., Veness, Joel, Bellemare, Marc G., Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K., Ostrovski, Georg, Petersen, Stig, Beattie, Charles, Sadik, Amir, Antonoglou, Ioannis, King, Helen, Kumaran, Dharshan, Wierstra, Daan, Legg, Shane, and Hassabis, Demis · 2015
Cited alongside, same era.
Universal value function approximators
Schaul, Tom, Horgan, Dan, Gregor, Karol, and Silver, David · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Bellemare, Marc, Srinivasan, Sriram, Ostrovski, Georg, Schaul, Tom, Saxton, David, and Munos, Remi
Cited in the paper.
Lake, Brenden M, Ullman, Tomer D, Tenenbaum, Joshua B, and Gershman, Samuel J · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Mnih, Volodymyr, Badia, Adria Puigdomenech, Mirza, Mehdi, Graves, Alex, Lillicrap, Timothy P, Harley, Tim, Silver, David, and Kavukcuoglu, Koray · 2016
Later among the works it cites.
High-dimensional continuous control using generalized advantage estimation
Schulman, John, Moritz, Philipp, Levine, Sergey, Jordan, Michael, and Abbeel, Pieter · 2016
Later among the works it cites.
A deep hierarchical approach to lifelong learning in minecraft
Tessler, Chen, Givony, Shahar, Zahavy, Tom, Mankowitz, Daniel J, and Mannor, Shie · 2016
Later among the works it cites.
Strategic attentive writer for learning macro-actions
Vezhnevets, Alexander, Mnih, Volodymyr, Osindero, Simon, Graves, Alex, Vinyals, Oriol, Agapiou, John, and kavukcuoglu, koray · 2016
Later among the works it cites.
Multi-scale context aggregation by dilated convolutions
Yu, Fisher and Koltun, Vladlen · 2016
Later among the works it cites.
The option-critic architecture
Bacon, Pierre-Luc, Precup, Doina, and Harb, Jean · 2017
Closest in time.