Fetching the paper…
Reading the bibliography…
Solving tasks with sparse rewards is a main challenge in reinforcement learning.
Feudal Reinforcement Learning
P. Dayan and G. Hinton · 1993
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
R. S. Sutton, D. Precup, and S. Singh · 1999
Earlier work this paper cites.
Bullet physics engine., 2009
E. Coumans · 2009
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
J. Schmidhuber · 2010
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
J. Kober, J. A. Bagnell, and J. Peters · 2013
Earlier work this paper cites.
Acquiring visual servoing reaching and grasping skills using neural reinforcement learning
T. Lampe and M. Riedmiller · 2013
Earlier work this paper cites.
End-to-End Training of Deep Visuomotor Policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
M. G. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos · 2016
Earlier work this paper cites.
Hierarchical relative entropy policy search
C. Daniel, G. Neumann, O. Kroemer, and J. Peters · 2016
Earlier work this paper cites.
Deep Reinforcement Learning for Robotic Manipulation
S. Gu, E. Holly, T. Lillicrap, and S. Levine · 2016
Earlier work this paper cites.
Learning and Transfer of Modulated Locomotor Controllers
N. Heess, G. Wayne, Y. Tassa, T. Lillicrap, M. Riedmiller, and D. Silver · 2016
Earlier work this paper cites.
Vime: Variational information maximizing exploration
R. Houthooft, X. Chen, Y. Duan, J. Schulman, F. De Turck, and P. Abbeel · 2016
Cited alongside, same era.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
T. D. Kulkarni, K. Narasimhan, A. Saeedi, and J. Tenenbaum · 2016
Cited alongside, same era.
Learning Hand-Eye Coordination for Robotic Grasping with Deep Learning and Large-Scale Data Collection
S. Levine, P. Pastor, A. Krizhevsky, and D. Quillen · 2016
Cited alongside, same era.
Supersizing self-supervision: Learning to grasp from 50K tries and 700 robot hours
L. Pinto and A. Gupta · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Later among the works it cites.
Data-efficient deep reinforcement learning for dexterous manipulation
I. Popov, N. Heess, T. Lillicrap, R. Hafner, G. Barth-Maron, M. Vecerik, T. Lampe, Y. Tassa, T. Erez, and M. Riedmiller · 2017
Later among the works it cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Later among the works it cites.
A deep hierarchical approach to lifelong learning in minecraft
C. Tessler, S. Givony, T. Zahavy, D. J. Mankowitz, and S. Mannor · 2017
Later among the works it cites.
FeUdal Networks for Hierarchical Reinforcement Learning
A. S. Vezhnevets, S. Osindero, T. Schaul, N. Heess, M. Jaderberg, D. Silver, and K. Kavukcuoglu · 2017
Later among the works it cites.
Meta learning shared hierarchies
K. Frans, J. Ho, X. Chen, P. Abbeel, and J. Schulman · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hindsight experience replay
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, P. Abbeel, and W. Zaremba · 2017
Cited alongside, same era.
The option-critic architecture
P.-L. Bacon, J. Harb, and D. Precup · 2017
Cited alongside, same era.
One-Shot Imitation Learning
Y. Duan, M. Andrychowicz, B. C. Stadie, J. Ho, J. Schneider, I. Sutskever, P. Abbeel, and W. Zaremba · 2017
Cited alongside, same era.
Stochastic Neural Networks for Hierarchical Reinforcement Learning
C. Florensa, Y. Duan, and P. Abbeel · 2017
Cited alongside, same era.
Tensorflow agents: Efficient batched reinforcement learning in tensorflow
D. Hafner, J. Davidson, and V. Vanhoucke · 2017
Cited alongside, same era.
Count-based exploration with neural density models
G. Ostrovski, M. G. Bellemare, A. van den Oord, and R. Munos · 2017
Cited alongside, same era.
Closest in time.
Latent space policies for hierarchical reinforcement learning
T. Haarnoja, K. Hartikainen, P. Abbeel, and S. Levine · 2018
Closest in time.
Learning an embedding space for transferable robot skills
K. Hausman, J. T. Springenberg, Z. Wang, N. Heess, and M. Riedmiller · 2018
Closest in time.
Deep reinforcement learning that matters
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger · 2018
Closest in time.
Data-efficient hierarchical reinforcement learning
O. Nachum, S. Gu, H. Lee, and S. Levine · 2018
Closest in time.
Diversity is all you need: Learning skills without a reward function
B. Eysenbach, A. Gupta, J. Ibarz, and S. Levine · 2019
Closest in time.