Fetching the paper…
Reading the bibliography…
Sparse-reward domains are challenging for reinforcement learning algorithms since significant exploration is needed before encountering reward for the first time.
Feudal reinforcement learning
Dayan, P. and Hinton, G. E · 1993
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the MAXQ value function decomposition
Dietterich, T. G · 2000
Earlier work this paper cites.
Dynamic abstraction in reinforcement learning via clustering
Mannor, S., Menache, I., Hoze, A., and Klein, U · 2004
Earlier work this paper cites.
An analysis of model-based interval estimation for Markov decision processes
Strehl, A. L. and Littman, M. L · 2008
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Earlier work this paper cites.
Safe Policy Search for Lifelong Reinforcement Learning with Sublinear Regret
Bou Ammar, H., R, T., and E, E · 2015
Earlier work this paper cites.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Cited alongside, same era.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, O. P., and Zaremba, W · 2017
Cited alongside, same era.
The option-critic architecture
Bacon, P.-L., Harb, J., and Precup, D · 2017
Cited alongside, same era.
Dilated Recurrent Neural Networks
Chang, S., Zhang, Y., Han, W., Yu, M., Guo, X., Tan, W., Cui, X., Witbrock, M., Hasegawa-Johnson, M., and Huang, T · 2017
Cited alongside, same era.
Automatic goal generation for reinforcement learning agents
Florensa, C., Held, D., Geng, X., and Abbeel, P · 2017
Cited alongside, same era.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Keramati, R., Whang, J., Cho, P., and Brunskill, E · 2018
Later among the works it cites.
Near-optimal representation learning for hierarchical reinforcement learning
Nachum, O., Gu, S., Lee, H., and Levine, S · 2018
Later among the works it cites.
Oh, J., Guo, Y., Singh, S., and Lee, H · 2018
Later among the works it cites.
Go-explore: a new approach for hard-exploration problems
Ecoffet, A., Huizinga, J., Lehman, J., Stanley, K. O., and Clune, J · 2019
Later among the works it cites.
Search on the replay buffer: Bridging planning and reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sutton, R. S., Modayil, J., Degris, M. D. T., Pilarski, P. M., and White, A · 2017
Cited alongside, same era.
Feudal networks for hierarchical reinforcement learning
Vezhnevets, A. S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D., and Kavukcuoglu, K · 2017
Cited alongside, same era.
Eysenbach, B., Salakhutdinov, R. R., and Levine, S · 2019
Later among the works it cites.
Learning World Graphs to Accelerate Hierarchical Reinforcement Learning
Shang, W., Trott, A., Sheng, S., Xiong, C., and Socher, R · 2019
Later among the works it cites.
Unsupervised Reinforcement Learning
Levine, S · 2020
Closest in time.