Fetching the paper…
Reading the bibliography…
Representation learning and option discovery are two of the biggest challenges in reinforcement learning (RL).
Dynamic Programming
Bellman, Richard E · 1957
Earlier work this paper cites.
Technical Note: Q-Learning
Watkins, Christopher J. C. H. and Dayan, Peter · 1992
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, Martin L · 1994
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, Richard S. and Barto, Andrew G · 1998
Earlier work this paper cites.
Between MDPs and semi-MDPs: A Framework for Temporal Abstraction in Reinforcement Learning
Sutton, Richard S., Precup, Doina, and Singh, Satinder · 1999
Earlier work this paper cites.
Hierarchical Reinforcement Learning with the MAXQ Value Function Decomposition
Dietterich, Thomas G · 2000
Earlier work this paper cites.
Temporal Abstraction in Reinforcement Learning
Precup, Doina · 2000
Earlier work this paper cites.
Automatic Discovery of Subgoals in Reinforcement Learning using Diverse Density
McGovern, Amy and Barto, Andrew G · 2001
Earlier work this paper cites.
Discovering Hierarchy in Reinforcement Learning with HEXQ
Hengst, Bernhard · 2002
Earlier work this paper cites.
Q-Cut - Dynamic Discovery of Sub-goals in Reinforcement Learning
Menache, Ishai, Mannor, Shie, and Shimkin, Nahum · 2002
Earlier work this paper cites.
Slow Feature Analysis: Unsupervised Learning of Invariances
Wiskott, Laurenz and Sejnowski, Terrence J · 2002
Earlier work this paper cites.
Using Relative Novelty to Identify Useful Temporal Abstractions in Reinforcement Learning
Şimşek, Özgür and Barto, Andrew G · 2004
Earlier work this paper cites.
Dynamic Abstraction in Reinforcement Learning via Clustering
Mannor, Shie, Menache, Ishai, Hoze, Amit, and Klein, Uri · 2004
Earlier work this paper cites.
Perron Cluster Analysis and Its Connection to Graph Partitioning for Noisy Data
Weber, Marcus, Rungsarityotin, Wasinee, and Schliep, Alexander · 2004
Earlier work this paper cites.
Identifying Useful Subgoals in Reinforcement Learning by Local Graph Partitioning
Şimşek, Özgür, Wolfe, Alicia P., and Barto, Andrew G · 2005
Cited alongside, same era.
Proto-Value Functions: Developmental Reinforcement Learning
Mahadevan, Sridhar · 2005
Cited alongside, same era.
Linear Algebra and Its Applications
Strang, Gilbert · 2005
Cited alongside, same era.
Graph Theory and Its Applications
Gross, Jonathan L. and Yellen, Jay · 2006
Cited alongside, same era.
Proto-value Functions: A Laplacian Framework for Learning Representation and Control in Markov Decision Processes
Mahadevan, Sridhar and Maggioni, Mauro · 2007
Cited alongside, same era.
Skill Characterization Based on Betweenness
Şimşek, Özgür and Barto, Andrew G · 2008
Cited alongside, same era.
Empowerment – An Introduction
Salge, Christoph, Glackin, Cornelius, and Polani, Daniel · 2014
Later among the works it cites.
Optimal Behavioral Hierarchy
Solway, Alec, Diuk, Carlos, Córdova, Natalia, Yee, Debbie, Barto, Andrew G., Niv, Yael, and Botvinick, Matthew M · 2014
Later among the works it cites.
Continual Curiosity-Driven Skill Acquisition from High-Dimensional Video Inputs for Humanoid Robots
Kompella, Varun Raj, Stollenga, Marijn, Luciw, Matthew, and Schmidhuber, Juergen · 2015
Later among the works it cites.
Probabilistic Inference for Determining Options in Reinforcement Learning
Daniel, Christian, van Hoof, Herke, Peters, Jan, and Neumann, Gerhard · 2016
Later among the works it cites.
Gregor, Karol, Rezende, Danilo, and Wierstra, Daan · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Skill Discovery in Continuous Reinforcement Learning Domains using Skill Chaining
Konidaris, George and Barto, Andrew · 2009
Cited alongside, same era.
Algorithms for Reinforcement Learning
Szepesvári, Csaba · 2010
Cited alongside, same era.
Incremental Slow Feature Analysis
Kompella, Varun Raj, Luciw, Matthew D., and Schmidhuber, Jürgen · 2011
Cited alongside, same era.
On the Relation of Slow Feature Analysis and Laplacian Eigenmaps
Sprekeler, Henning · 2011
Cited alongside, same era.
On the Bottleneck Concept for Options Discovery: Theoretical Underpinnings and Extension in Continuous State Spaces
Bacon, Pierre-Luc · 2013
Cited alongside, same era.
Active learning of inverse models with intrinsically motivated goal exploration in robots
Baranes, Adrien and Oudeyer, Pierre-Yves · 2013
Cited alongside, same era.
Kulkarni, Tejas D., Narasimhan, Karthik R., Saeedi, Ardavan, and Tenenbaum, Joshua B · 2016
Later among the works it cites.
Option Discovery in Hierarchical Reinforcement Learning using Spatio-Temporal Clustering
Lakshminarayanan, Aravind, Krishnamurthy, Ramnandan, Kumar, Peeyush, and Ravindran, Balaraman · 2016
Later among the works it cites.
Learning Purposeful Behaviour in the Absence of Rewards
Machado, Marlos C. and Bowling, Michael · 2016
Later among the works it cites.
Adaptive Skills Adaptive Partitions (ASAP)
Mankowitz, Daniel J., Mann, Timothy Arthur, and Mannor, Shie · 2016
Later among the works it cites.
Control of Memory, Active Perception, and Action in Minecraft
Oh, Junhyuk, Chockalingam, Valliappa, Singh, Satinder P., and Lee, Honglak · 2016
Later among the works it cites.
Generalization and Exploration via Randomized Value Functions
Osband, Ian, Roy, Benjamin Van, and Wen, Zheng · 2016
Later among the works it cites.
Strategic Attentive Writer for Learning Macro-Actions
Vezhnevets, Alexander, Mnih, Volodymyr, Osindero, Simon, Graves, Alex, Vinyals, Oriol, Agapiou, John, and Kavukcuoglu, Koray · 2016
Later among the works it cites.
The option-critic architecture
Bacon, Pierre-Luc, Harb, Jean, and Precup, Doina · 2017
Closest in time.