Fetching the paper…
Reading the bibliography…
Learning goal-directed behavior in environments with sparse feedback is a major challenge for reinforcement learning algorithms.
Improving generalization for temporal difference learning: The successor representation
P. Dayan · 1993
Earlier work this paper cites.
Feudal reinforcement learning
P. Dayan and G. E. Hinton · 1993
Earlier work this paper cites.
Introduction to reinforcement learning
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
R. S. Sutton, D. Precup, and S. Singh · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
T. G. Dietterich · 2000
Earlier work this paper cites.
Hierarchical memory-based reinforcement learning
N. Hernandez-Gardiol and S. Mahadevan · 2001
Earlier work this paper cites.
Automatic discovery of subgoals in reinforcement learning using diverse density
A. McGovern and A. G. Barto · 2001
Earlier work this paper cites.
Q-cut—dynamic discovery of sub-goals in reinforcement learning
I. Menache, S. Mannor, and N. Shimkin · 2002
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
A. G. Barto and S. Mahadevan · 2003
Earlier work this paper cites.
Subgoal discovery for hierarchical reinforcement learning using learned policies
S. Goel and M. Huber · 2003
Earlier work this paper cites.
Generalizing plans to new environments in relational mdps
C. Guestrin, D. Koller, C. Gearhart, and N. Kanodia · 2003
Earlier work this paper cites.
Efficient solution algorithms for factored mdps
C. Guestrin, D. Koller, R. Parr, and S. Venkataraman · 2003
Earlier work this paper cites.
Dynamic abstraction in reinforcement learning via clustering
S. Mannor, I. Menache, A. Hoze, and U. Klein · 2004
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
S. P. Singh, A. G. Barto, and N. Chentanez · 2004
Earlier work this paper cites.
Identifying useful subgoals in reinforcement learning by local graph partitioning
Ö. Şimşek, A. Wolfe, and A. Barto · 2005
Earlier work this paper cites.
Core knowledge
E. S. Spelke and K. D. Kinzler · 2007
Earlier work this paper cites.
An object-oriented representation for efficient reinforcement learning
C. Diuk, A. Cohen, and M. L. Littman · 2008
Earlier work this paper cites.
Hierarchically organized behavior and its neural foundations: A reinforcement learning perspective
M. M. Botvinick, Y. Niv, and A. C. Barto · 2009
Cited alongside, same era.
Where do rewards come from
S. Singh, R. L. Lewis, and A. G. Barto · 2009
Cited alongside, same era.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
J. Schmidhuber · 2010
Cited alongside, same era.
Intrinsically motivated reinforcement learning: An evolutionary perspective
S. Singh, R. L. Lewis, A. G. Barto, and J. Sorg · 2010
Cited alongside, same era.
Linear options
J. Sorg and S. Singh · 2010
Cited alongside, same era.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
R. S. Sutton, J. Modayil, M. Delp, T. Degris, P. M. Pilarski, A. White, and D. Precup · 2011
Efficient inference in occlusion-aware generative models of images
J. Huang and K. Murphy · 2015
Later among the works it cites.
Deep convolutional inverse graphics network
T. D. Kulkarni, W. F. Whitney, P. Kohli, and J. Tenenbaum · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Later among the works it cites.
Variational information maximisation for intrinsically motivated reinforcement learning
S. Mohamed and D. J. Rezende · 2015
Later among the works it cites.
Massively parallel methods for deep reinforcement learning
A. Nair, P. Srinivasan, S. Blackwell, C. Alcicek, R. Fearon, A. De Maria, V. Panneershelvam, M. Suleyman, C. Beattie, S. Petersen, et al · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2012
Cited alongside, same era.
The successor representation and temporal context
S. J. Gershman, C. D. Moore, M. T. Todd, K. A. Norman, and P. B. Sederberg · 2012
Cited alongside, same era.
Object focused q-learning for autonomous agents
L. C. Cobo, C. L. Isbell, and A. L. Thomaz · 2013
Cited alongside, same era.
Evolving deep unsupervised convolutional networks for vision-based reinforcement learning
J. Koutník, J. Schmidhuber, and F. Gomez · 2014
Cited alongside, same era.
Design principles of the hippocampal cognitive map
K. L. Stachenfeld, M. Botvinick, and S. J. Gershman · 2014
Cited alongside, same era.
Universal option models
C. Szepesvari, R. S. Sutton, J. Modayil, S. Bhatnagar, et al · 2014
Cited alongside, same era.
Language understanding for text-based games using deep reinforcement learning
K. Narasimhan, T. Kulkarni, and R. Barzilay · 2015
Later among the works it cites.
Universal value function approximators
T. Schaul, D. Horgan, K. Gregor, and D. Silver · 2015
Later among the works it cites.
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2015
Later among the works it cites.
Incentivizing exploration in reinforcement learning with deep predictive models
B. C. Stadie, S. Levine, and P. Abbeel · 2015
Later among the works it cites.
Attend, infer, repeat: Fast scene understanding with generative models
S. Eslami, N. Heess, T. Weber, Y. Tassa, K. Kavukcuoglu, and G. E. Hinton · 2016
Closest in time.
Building machines that learn and think like people
B. M. Lake, T. D. Ullman, J. B. Tenenbaum, and S. J. Gershman · 2016
Closest in time.
Learning purposeful behaviour in the absence of rewards
M. C. Machado and M. Bowling · 2016
Closest in time.
Deep exploration via bootstrapped dqn
I. Osband, C. Blundell, A. Pritzel, and B. Van Roy · 2016
Closest in time.
One-shot generalization in deep generative models
D. J. Rezende, S. Mohamed, I. Danihelka, K. Gregor, and D. Wierstra · 2016
Closest in time.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Closest in time.
Understanding visual concepts with continuation learning
W. F. Whitney, M. Chang, T. Kulkarni, and J. B. Tenenbaum · 2016
Closest in time.