Fetching the paper…
Reading the bibliography…
In hierarchical reinforcement learning a major challenge is determining appropriate low-level policies.
Curious model-building control systems
J. Schmidhuber · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
R. S. Sutton, D. Precup, and S. P. Singh · 1999
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
S. P. Singh, A. G. Barto, and N. Chentanez · 2004
Earlier work this paper cites.
Empowerment: a universal agent-centric measure of control
A. S. Klyubin, D. Polani, and C. L. Nehaniv · 2005
Earlier work this paper cites.
What is intrinsic motivation? A typology of computational approaches
P. Oudeyer and F. Kaplan · 2009
Earlier work this paper cites.
Lecture 6.5 - rmsprop, coursera: Neural networks for machine learning, 2012
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Cited alongside, same era.
Variational information maximisation for intrinsically motivated reinforcement learning
S. Mohamed and D. J. Rezende · 2015
Cited alongside, same era.
Mazebase: A sandbox for learning from games
S. Sukhbaatar, A. Szlam, G. Synnaeve, S. Chintala, and R. Fergus · 2015
Cited alongside, same era.
Benchmarking deep reinforcement learning for continuous control
Y. Duan, X. Chen, R. Houthooft, J. Schulman, and P. Abbeel · 2016
Cited alongside, same era.
K. Gregor, D. J. Rezende, and D. Wierstra · 2016
Cited alongside, same era.
The option-critic architecture
P.-L. Bacon, J. Harb, and D. Precup · 2017
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Later among the works it cites.
Feudal networks for hierarchical reinforcement learning
A. S. Vezhnevets, S. Osindero, T. Schaul, N. Heess, M. Jaderberg, D. Silver, and K. Kavukcuoglu · 2017
Later among the works it cites.
Diversity is all you need: Learning skills without a reward function
B. Eysenbach, A. Gupta, J. Ibarz, and S. Levine · 2018
Closest in time.
Latent space policies for hierarchical reinforcement learning
T. Haarnoja, K. Hartikainen, P. Abbeel, and S. Levine · 2018
Closest in time.
Learning an embedding space for transferable robot skills
K. Hausman, J. T. Springenberg, Z. Wang, N. Heess, and M. Riedmiller · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Reverse curriculum generation for reinforcement learning
C. Florensa, D. Held, M. Wulfmeier, M. Zhang, and P. Abbeel · 2017
Cited alongside, same era.
Unsupervised learning of goal spaces for intrinsically motivated goal exploration
A. Péré, S. Forestier, O. Sigaud, and P.-Y. Oudeyer · 2018
Closest in time.
Intrinsic motivation and automatic curricula via asymmetric self-play
S. Sukhbaatar, Z. Lin, I. Kostrikov, G. Synnaeve, A. Szlam, and R. Fergus · 2018
Closest in time.