Fetching the paper…
Reading the bibliography…
In this work, we consider the problem of autonomously discovering behavioral abstractions, or options, for reinforcement learning agents.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K. (2016) · 1937
Earlier work this paper cites.
Dynamic programming
Bellman, R. (1957) · 1957
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J. (1992) · 1992
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L. (1994) · 1994
Earlier work this paper cites.
Stochastic approximation with two time scales
Borkar, V. S. (1997) · 1997
Earlier work this paper cites.
Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S. (1999) · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y. (2000) · 2000
Earlier work this paper cites.
Automatic discovery of subgoals in reinforcement learning using diverse density
McGovern, A. and Barto, A. G. (2001) · 2001
Earlier work this paper cites.
Learning options in reinforcement learning
Stolle, M. and Precup, D. (2003) · 2003
Earlier work this paper cites.
Optimal behavioral hierarchy
Solway, A., Diuk, C., Córdova, N., Yee, D., Barto, A. G., Niv, Y., and Botvinick, M. M. (2014) · 2014
Cited alongside, same era.
Probabilistic segmentation applied to an assembly task
Lioutikov, R., Neumann, G., Maeda, G., and Peters, J. (2015) · 2015
Cited alongside, same era.
Probabilistic inference for determining options in reinforcement learning
Daniel, C., van Hoof, H., Peters, J., and Neumann, G. (2016) · 2016
Cited alongside, same era.
Gregor, K., Rezende, D. J., and Wierstra, D. (2016) · 2016
Cited alongside, same era.
Learning and transfer of modulated locomotor controllers
Heess, N., Wayne, G., Tassa, Y., Lillicrap, T. P., Riedmiller, M. A., and Silver, D. (2016) · 2016
Cited alongside, same era.
Ddco: Discovery of deep continuous options for robot learning from demonstrations
Krishnan, S., Fox, R., Stoica, I., and Goldberg, K. (2017) · 2017
Later among the works it cites.
Reinforcement Learning: An Introduction
Sutton, R. S. and Barto, A. G. (2017) · 2017
Later among the works it cites.
Feudal networks for hierarchical reinforcement learning
Vezhnevets, A. S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D., and Kavukcuoglu, K. (2017) · 2017
Later among the works it cites.
Self-Consistent Trajectory Autoencoder: Hierarchical Reinforcement Learning with Trajectory Embeddings
Co-Reyes, J. D., Liu, Y., Gupta, A., Eysenbach, B., Abbeel, P., and Levine, S. (2018) · 2018
Later among the works it cites.
Meta learning shared hierarchies
Frans, K., Ho, J., Chen, X., Abbeel, P., and Schulman, J. (2018) · 2018
Later among the works it cites.
When waiting is not an option: Learning options with a deliberation cost
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The option-critic architecture
Bacon, P.-L., Harb, J., and Precup, D. (2017) · 2017
Cited alongside, same era.
Stochastic neural networks for hierarchical reinforcement learning
Florensa, C., Duan, Y., and Abbeel, P. (2017) · 2017
Cited alongside, same era.
Multi-level discovery of deep options
Fox, R., Krishnan, S., Stoica, I., and Goldberg, K. (2017) · 2017
Cited alongside, same era.
Harb, J., Bacon, P.-L., Klissarov, M., and Precup, D. (2018) · 2018
Later among the works it cites.
Learning an embedding space for transferable robot skills
Hausman, K., Springenberg, J. T., Wang, Z., Heess, N., and Riedmiller, M. (2018) · 2018
Later among the works it cites.
Diversity is all you need: Learning skills without a reward function
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S. (2019) · 2019
Closest in time.