How to combine tree-search methods in reinforcement learning
Original
Efroni, Y., Dalal, G., Scherrer, B., and Mannor, S · 2018
Later among the works it cites.
IMPALA: Scalable distributed deep-RL with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., Legg, S., and Kavukcuoglu, K · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Original
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Later among the works it cites.
Learning an embedding space for transferable robot skills
Hausman, K., Springenberg, J. T., Wang, Z., Heess, N., and Riedmiller, M · 2018
Later among the works it cites.
Plan online, learn offline: Efficient learning and exploration via model-based control
Original
Lowrey, K., Rajeswaran, A., Kakade, S. M., Todorov, E., and Mordatch, I · 2018
Later among the works it cites.
A0c: Alpha zero in continuous action space, 2018
Moerland, T. M., Broekens, J., Plaat, A., and Jonker, C. M · 2018
Later among the works it cites.
Deepmimic: Example-guided deep reinforcement learning of physics-based character skills
Peng, X. B., Abbeel, P., Levine, S., and van de Panne, M · 2018
Later among the works it cites.
DeepMind control suite
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., de Las Casas, D., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., Lillicrap, T., and Riedmiller, M · 2018
Later among the works it cites.
Policy gradient search: Online planning and expert iteration without search trees
Original
Anthony, T., Nishihara, R., Moritz, P., Salimans, T., and Schulman, J · 2019
Later among the works it cites.
Information theoretic model predictive q-learning, 2019
Bhardwaj, M., Handa, A., Fox, D., and Boots, B · 2019
Later among the works it cites.
Approximate inference in discrete distributions with monte carlo tree search and value functions, 2019
Buesing, L., Heess, N., and Weber, T · 2019
Later among the works it cites.
Imagined value gradients: Model-based policy optimization with transferable latent dynamics models
Original
Byravan, A., Springenberg, J. T., Abdolmaleki, A., Hafner, R., Neunert, M., Lampe, T., Siegel, N., Heess, N., and Riedmiller, M · 2019
Later among the works it cites.
Monte-carlo tree search for policy optimization, 2019
Ma, X., Driggs-Campbell, K., Zhang, Z., and Kochenderfer, M. J · 2019
Later among the works it cites.
Hierarchical visuomotor control of humanoids
Merel, J., Ahuja, A., Pham, V., Tunyasuvunakool, S., Liu, S., Tirumala, D., Heess, N., and Wayne, G · 2019
Later among the works it cites.
Probabilistic planning with sequential monte carlo methods
Piché, A., Thomas, V., Ibrahim, C., Bengio, Y., and Pal, C · 2019
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model, 2019
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., Lillicrap, T., and Silver, D · 2019
Later among the works it cites.
V-mpo: On-policy maximum a posteriori policy optimization for discrete and continuous control, 2019
Song, H. F., Abdolmaleki, A., Springenberg, J. T., Clark, A., Soyer, H., Rae, J. W., Noury, S., Ahuja, A., Liu, S., Tirumala, D., Heess, N., Belov, D., Riedmiller, M., and Botvinick, M. M · 2019
Later among the works it cites.
When to use parametric models in reinforcement learning?
Original
van Hasselt, H., Hessel, M., and Aslanides, J · 2019
Later among the works it cites.