Fetching the paper…
Reading the bibliography…
Recent work has shown that temporally extended actions (options) can be learned fully end-to-end as opposed to being specified in advance.
Models of man: social and rational; mathematical essays on rational human behavior in society setting
Herbert A. Simon · 1957
Earlier work this paper cites.
Steps toward artificial intelligence
Marvin Minsky · 1961
Earlier work this paper cites.
Semi-markovian decision processes
Ronald A. Howard · 1963
Earlier work this paper cites.
Learning and executing generalized robot plans
Richard Fikes, Peter E. Hart, and Nils J. Nilsson · 1972
Earlier work this paper cites.
Commonsense knowledge of space: Learning from experience
Benjamin Kuipers · 1979
Earlier work this paper cites.
Learning to Solve Problems by Searching for Macro-operators
Richard Earl Korf · 1983
Earlier work this paper cites.
Temporal credit assignment in reinforcement learning
Richard S. Sutton · 1984
Earlier work this paper cites.
Bounded complexity justifies cooperation in the finitely repeated prisoners dilemma
Abraham Neyman · 1985
Earlier work this paper cites.
A heuristic approach to the discovery of macro-operators
Glenn A. Iba · 1989
Earlier work this paper cites.
Made-up Minds: A Constructivist Approach to Artificial Intelligence
Gary L. Drescher · 1991
Earlier work this paper cites.
Constrained discounted markov decision chains
Linn I. Sennott · 1991
Earlier work this paper cites.
Feudal reinforcement learning
Peter Dayan and Geoffrey E. Hinton · 1992
Earlier work this paper cites.
Advantage updating
Leemon C. Baird · 1993
Earlier work this paper cites.
Hierarchical learning in stochastic domains: Preliminary results
Leslie Pack Kaelbling · 1993
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Cited alongside, same era.
Finding structure in reinforcement learning
Sebastian Thrun and Anton Schwartz · 1995
Cited alongside, same era.
The MAXQ method for hierarchical reinforcement learning
Thomas G. Dietterich · 1998
Cited alongside, same era.
Reinforcement learning with hierarchies of machines
Ronald Parr and Stuart J. Russell · 1998
Cited alongside, same era.
Constrained Markov Decision Processes
E. Altman · 1999
Cited alongside, same era.
Policy gradient methods for reinforcement learning with function approximation
Richard S. Sutton, David A. McAllester, Satinder P. Singh, and Yishay Mansour · 1999
Cited alongside, same era.
Deep learning for real-time atari game play using offline monte-carlo tree search planning
Xiaoxiao Guo, Satinder Singh, Honglak Lee, Richard L Lewis, and Xiaoshi Wang · 2014
Later among the works it cites.
Time-regularized interrupting options (TRIO)
Timothy Arthur Mann, Daniel J. Mankowitz, and Shie Mannor · 2014
Later among the works it cites.
Optimal behavioral hierarchy
Alec Solway, Carlos Diuk, Natalia Córdova, Debbie Yee, Andrew G. Barto, Yael Niv, and Matthew M. Botvinick · 2014
Later among the works it cites.
The dependence of effective planning horizon on model accuracy
Nan Jiang, Alex Kulesza, Satinder Singh, and Richard L. Lewis · 2015
Later among the works it cites.
Approximate value iteration with temporally extended actions
Timothy Arthur Mann, Shie Mannor, and Doina Precup · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Richard S. Sutton, Doina Precup, and Satinder P. Singh · 1999
Cited alongside, same era.
Temporal abstraction in reinforcement learning
Doina Precup · 2000
Cited alongside, same era.
Bounded Rationality: The adaptive toolbox
Gerd Gigerenzer and R. Selten · 2001
Cited alongside, same era.
Biasing approximate dynamic programming with a lower discount factor
Marek Petrik and Bruno Scherrer · 2008
Cited alongside, same era.
Hierarchically organized behavior and its neural foundations: A reinforcement learning perspective
Matthew M. Botvinick, Yael Niv, and Andrew C. Barto · 2009
Cited alongside, same era.
Learning high-level planning from text
S. R. K. Branavan, Nate Kushman, Tao Lei, and Regina Barzilay · 2012
Cited alongside, same era.
Later among the works it cites.
Probabilistic inference for determining options in reinforcement learning
C. Daniel, H. van Hoof, J. Peters, and G. Neumann · 2016
Later among the works it cites.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Tejas Kulkarni, Karthik Narasimhan, Ardavan Saeedi, and Joshua Tenenbaum · 2016
Later among the works it cites.
Adaptive skills, adaptive partitions (ASAP)
Daniel J. Mankowitz, Timothy Arthur Mann, and Shie Mannor · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Later among the works it cites.
Modular multitask reinforcement learning with policy sketches
Jacob Andreas, Dan Klein, and Sergey Levine · 2017
Closest in time.
The option-critic architecture
Pierre-Luc Bacon, Jean Harb, and Doina Precup · 2017
Closest in time.
A laplacian framework for option discovery in reinforcement learning
Marlos C. Machado, Marc G. Bellemare, and Michael H. Bowling · 2017
Closest in time.