Fetching the paper…
Reading the bibliography…
The problem of sparse rewards is one of the hardest challenges in contemporary reinforcement learning.
Curious model-building control systems
J. Schmidhuber · 1991
Earlier work this paper cites.
Feudal reinforcement learning
P. Dayan and G. E. Hinton · 1992
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Reinforcement learning with hierarchies of machines
R. Parr and S. J. Russell · 1998
Earlier work this paper cites.
Introduction to Reinforcement Learning
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
R. S. Sutton, D. Precup, and S. Singh · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
T. G. Dietterich · 2000
Earlier work this paper cites.
Automatic discovery of subgoals in reinforcement learning using diverse density
A. McGovern and A. G. Barto · 2001
Earlier work this paper cites.
Q-cut—dynamic discovery of sub-goals in reinforcement learning
I. Menache, S. Mannor, and N. Shimkin · 2002
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
A. G. Barto and S. Mahadevan · 2003
Earlier work this paper cites.
Using relative novelty to identify useful temporal abstractions in reinforcement learning
Ö. Şimşek and A. G. Barto · 2004
Earlier work this paper cites.
Empowerment: A universal agent-centric measure of control
A. S. Klyubin, D. Polani, and C. L. Nehaniv · 2005
Cited alongside, same era.
Skill discovery in continuous reinforcement learning domains using skill chaining
G. Konidaris and A. G. Barto · 2009
Cited alongside, same era.
The arcade learning environment: An evaluation platform for general agents
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling · 2013
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mané, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Viégas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2015
Cited alongside, same era.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
T. D. Kulkarni, K. Narasimhan, A. Saeedi, and J. Tenenbaum · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Later among the works it cites.
High-dimensional continuous control using generalized advantage estimation
J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel · 2016
Later among the works it cites.
Learning to reinforcement learn
J. X. Wang, Z. Kurth-Nelson, D. Tirumala, H. Soyer, J. Z. Leibo, R. Munos, C. Blundell, D. Kumaran, and M. Botvinick · 2016
Later among the works it cites.
The option-critic architecture
P.-L. Bacon, J. Harb, and D. Precup · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
Variational information maximisation for intrinsically motivated reinforcement learning
S. Mohamed and D. J. Rezende · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
M. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos · 2016
Cited alongside, same era.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Cited alongside, same era.
Vime: Variational information maximizing exploration
R. Houthooft, X. Chen, Y. Duan, J. Schulman, F. De Turck, and P. Abbeel · 2016
Cited alongside, same era.
E. Bengio, V. Thomas, J. Pineau, D. Precup, and Y. Bengio · 2017
Closest in time.
Variational Intrinsic Control
K. Gregor, D. Jimenez Rezende, and D. Wierstra · 2017
Closest in time.
Reinforcement learning with unsupervised auxiliary tasks
M. Jaderberg, V. Mnih, W. M. Czarnecki, T. Schaul, J. Z. Leibo, D. Silver, and K. Kavukcuoglu · 2017
Closest in time.
Count-based exploration with neural density models
G. Ostrovski, M. G. Bellemare, A. van den Oord, and R. Munos · 2017
Closest in time.
Feudal networks for hierarchical reinforcement learning
A. S. Vezhnevets, S. Osindero, T. Schaul, N. Heess, M. Jaderberg, D. Silver, and K. Kavukcuoglu · 2017
Closest in time.