Fetching the paper…
Reading the bibliography…
We reformulate the option framework as two parallel augmented MDPs.
Feudal reinforcement learning
Dayan, P. and Hinton, G. E. (1993) · 1993
Earlier work this paper cites.
Planning simple trajectories using neural subgoal generators
Schmidhuber, J. and Wahnsiedler, R. (1993) · 1993
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S. (1999) · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
Dietterich, T. G. (2000) · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y. (2000) · 2000
Earlier work this paper cites.
Automatic discovery of subgoals in reinforcement learning using diverse density
McGovern, A. and Barto, A. G. (2001) · 2001
Earlier work this paper cites.
Learning options in reinforcement learning
Stolle, M. and Precup, D. (2002) · 2002
Earlier work this paper cites.
Natural actor-critic
Peters, J. and Schaal, S. (2008) · 2008
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Nair, V. and Hinton, G. E. (2010) · 2010
Earlier work this paper cites.
Unified inter and intra options learning using policy gradient methods
Levy, K. Y. and Shimkin, N. (2011) · 2011
Earlier work this paper cites.
Clustering via dirichlet process mixture models for portable skill discovery
Niekum, S. and Barto, A. G. (2011) · 2011
Earlier work this paper cites.
Compositional planning using optimal option models
Silver, D. and Ciosek, K. (2012) · 2012
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L. (2014) · 2014
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P. (2015) · 2015
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Cited alongside, same era.
Probabilistic inference for determining options in reinforcement learning
Daniel, C., Van Hoof, H., Peters, J., and Neumann, G. (2016) · 2016
Cited alongside, same era.
The option-critic architecture
Bacon, P.-L., Harb, J., and Precup, D. (2017) · 2017
Cited alongside, same era.
Feudal networks for hierarchical reinforcement learning
Vezhnevets, A. S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D., and Kavukcuoglu, K. (2017) · 2017
Later among the works it cites.
Ace: An actor ensemble algorithm for continuous control with tree search
Zhang, S., Chen, H., and Yao, H. (2019b) · 2017
Later among the works it cites.
Temporal Representation Learning
Bacon, P.-L. (2018) · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. (2018) · 2018
Later among the works it cites.
Learning latent dynamics for planning from pixels
Hafner, D., Lillicrap, T., Fischer, I., Villegas, R., Ha, D., Lee, H., and Davidson, J. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ciosek, K. and Whiteson, S. (2017) · 2017
Cited alongside, same era.
Efficient parallel methods for deep reinforcement learning
Clemente, A. V., Castejón, H. N., and Chandra, A. (2017) · 2017
Cited alongside, same era.
Openai baselines
Dhariwal, P., Hesse, C., Klimov, O., Nichol, A., Plappert, M., Radford, A., Schulman, J., Sidor, S., Wu, Y., and Zhokhov, P. (2017) · 2017
Cited alongside, same era.
Stochastic neural networks for hierarchical reinforcement learning
Florensa, C., Duan, Y., and Abbeel, P. (2017) · 2017
Cited alongside, same era.
Learnings options end-to-end for continuous action tasks
Klissarov, M., Bacon, P.-L., Harb, J., and Precup, D. (2017) · 2017
Cited alongside, same era.
Levy, A., Platt, R., and Saenko, K. (2017) · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Cited alongside, same era.
When waiting is not an option: Learning options with a deliberation cost
Harb, J., Bacon, P.-L., Klissarov, M., and Precup, D. (2018) · 2018
Later among the works it cites.
Data-efficient hierarchical reinforcement learning
Nachum, O., Gu, S. S., Lee, H., and Levine, S. (2018) · 2018
Later among the works it cites.
Openai five
OpenAI (2018) · 2018
Later among the works it cites.
Temporal abstraction
Precup, D. (2018) · 2018
Later among the works it cites.
Learning abstract options
Riemer, M., Liu, M., and Tesauro, G. (2018) · 2018
Later among the works it cites.
An inference-based policy gradient method for learning options
Smith, M., Hoof, H., and Pineau, J. (2018) · 2018
Later among the works it cites.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. d. L., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., et al. (2018) · 2018
Later among the works it cites.