Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D · 2016
Later among the works it cites.
The option-critic architecture
Bacon, P.-L., Harb, J., and Precup, D · 2017
Later among the works it cites.
Stochastic neural networks for hierarchical reinforcement learning
Florensa, C., Duan, Y., and P., A · 2017
Later among the works it cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Later among the works it cites.
Bridging the gap between value and policy based reinforcement learning
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D · 2017
Later among the works it cites.
Hierarchy through composition with multitask LMDPs
Saxe, A. M., Earle, A. C., and Rosman, B. S · 2017
Later among the works it cites.
A deep hierarchical approach to lifelong learning in minecraft
Tessler, C., Givony, S., Zahavy, T., Mankowitz, D. J., and Mannor, S · 2017
Later among the works it cites.
Feudal networks for hierarchical reinforcement learning
Original
Vezhnevets, A. S., Osindero, S., Schaul, T., Heess, N., Jaderberg, M., Silver, D., and Kavukcuoglu, K · 2017
Later among the works it cites.
Diversity is all you need: Learning skills without a reward function
Original
Eysenbach, B., Gupta, A., Ibarz, J., and Levine, S · 2018
Closest in time.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Original
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Closest in time.
Learning an embedding space for transferable robot skills
Hausman, K., Springenberg, J. T., Wang, Z., Heess, N., and Riedmiller, M · 2018
Closest in time.
Overlapping layered learning
MacAlpine, P. and Stone, P · 2018
Closest in time.