Fetching the paper…
Reading the bibliography…
We introduce Hindsight Off-policy Options (HO2), a data-efficient option learning algorithm.
Harutyunyan, A., Dabney, W., Borsa, D., Heess, N., Munos, R., and Precup, D · 1902
Earlier work this paper cites.
A tutorial on hidden markov models and selected applications in speech recognition
Rabiner, L. R · 1989
Earlier work this paper cites.
Feudal reinforcement learning
Dayan, P. and Hinton, G. E · 1993
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
Temporal abstraction in reinforcement learning
Precup, D · 2000
Earlier work this paper cites.
Off-policy learning with options and recognizers
Precup, D., Paduraru, C., Koop, A., Sutton, R. S., and Singh, S. P · 2006
Earlier work this paper cites.
Reinforcement learning with a gaussian mixture model
Agostini, A. and Celaya, E · 2010
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Moritz, P., Jordan, M., and Abbeel, P · 2015
Earlier work this paper cites.
Tensorflow: A system for large-scale machine learning
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., et al · 2016
Earlier work this paper cites.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Hierarchical relative entropy policy search
Daniel, C., Neumann, G., Kroemer, O., and Peters, J · 2016
Earlier work this paper cites.
Learning and transfer of modulated locomotor controllers
Heess, N., Wayne, G., Tassa, Y., Lillicrap, T., Riedmiller, M., and Silver, D · 2016
Earlier work this paper cites.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, P., and Zaremba, W · 2017
Earlier work this paper cites.
The option-critic architecture
Bacon, P.-L., Harb, J., and Precup, D · 2017
Cited alongside, same era.
The intentional unintentional agent: Learning to solve many continuous control tasks simultaneously
Cabi, S., Colmenarejo, S. G., Hoffman, M. W., Denil, M., Wang, Z., and Freitas, N · 2017
Cited alongside, same era.
Multi-level discovery of deep options
Fox, R., Krishnan, S., Stoica, I., and Goldberg, K · 2017
Cited alongside, same era.
Ddco: Discovery of deep continuous options for robot learning from demonstrations
Krishnan, S., Fox, R., Stoica, I., and Goldberg, K · 2017
Cited alongside, same era.
Learning multi-level hierarchies with hindsight
Levy, A., Konidaris, G., Platt, R., and Saenko, K · 2017
Cited alongside, same era.
When waiting is not an option: Learning options with a deliberation cost
Harb, J., Bacon, P.-L., Klissarov, M., and Precup, D · 2018
Later among the works it cites.
Learning an embedding space for transferable robot skills
Hausman, K., Springenberg, J. T., Wang, Z., Heess, N., and Riedmiller, M · 2018
Later among the works it cites.
Learning dexterous in-hand manipulation
OpenAI, Andrychowicz, M., Baker, B., Chociej, M., Józefowicz, R., McGrew, B., Pachocki, J. W., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., Schneider, J., Sidor, S., Tobin, J., Welinder, P., Weng, L., and Zaremba, W · 2018
Later among the works it cites.
Learning by playing-solving sparse reward tasks from scratch
Riedmiller, M., Hafner, R., Lampe, T., Neunert, M., Degrave, J., Van de Wiele, T., Mnih, V., Heess, N., and Springenberg, J. T · 2018
Later among the works it cites.
Learning abstract options
Riemer, M., Liu, M., and Tesauro, G · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Asymmetric actor critic for image-based robot learning
Pinto, L., Andrychowicz, M., Welinder, P., Zaremba, W., and Abbeel, P · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Cited alongside, same era.
Distral: Robust multitask reinforcement learning
Teh, Y. W., Bapst, V., Czarnecki, W. M., Quan, J., Kirkpatrick, J., Hadsell, R., Heess, N., and Pascanu, R · 2017
Cited alongside, same era.
Feudal networks for hierarchical reinforcement learning
Vezhnevets, A. S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D., and Kavukcuoglu, K · 2017
Cited alongside, same era.
Meta learning shared hierarchies
Frans, K., Ho, J., Chen, X., Abbeel, P., and Schulman, J · 2018
Cited alongside, same era.
Information asymmetry in kl-regularized rl
Galashov, A., Jayakumar, S. M., Hasenclever, L., Tirumala, D., Schwarz, J., Desjardins, G., Czarnecki, W. M., Teh, Y. W., Pascanu, R., and Heess, N · 2018
Cited alongside, same era.
Taco: Learning task decomposition via temporal alignment for control
Shiarlis, K., Wulfmeier, M., Salter, S., Whiteson, S., and Posner, I · 2018
Later among the works it cites.
An inference-based policy gradient method for learning options
Smith, M., Hoof, H., and Pineau, J · 2018
Later among the works it cites.
Multitask soft option learning
Igl, M., Gambardella, A., Nardelli, N., Siddharth, N., Böhmer, W., and Whiteson, S · 2019
Later among the works it cites.
Sub-policy adaptation for hierarchical reinforcement learning
Li, A. C., Florensa, C., Clavera, I., and Abbeel, P · 2019
Later among the works it cites.
Exploiting hierarchy for learning and transfer in kl-regularized rl
Tirumala, D., Noh, H., Galashov, A., Hasenclever, L., Ahuja, A., Wayne, G., Pascanu, R., Teh, Y. W., and Heess, N · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Later among the works it cites.
Dac: The double actor-critic architecture for learning options
Zhang, S. and Whiteson, S · 2019
Later among the works it cites.
Compositional Transfer in Hierarchical Reinforcement Learning
Wulfmeier, M., Abdolmaleki, A., Hafner, R., Springenberg, J. T., Neunert, M., Siegel, N., Hertweck, T., Lampe, T., Heess, N., and Riedmiller, M · 2020
Closest in time.