Deep exploration via bootstrapped dqn
Original
Osband, I., Blundell, C., Pritzel, A., and Roy, B. V. (2016) · 2016
Later among the works it cites.
Prioritized experience replay
Schaul, T., Quan, J., Antonoglou, I., and Silver, D. (2016) · 2016
Later among the works it cites.
Mastering the game of go using deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D. (2016) · 2016
Later among the works it cites.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., van Hasselt, H., Lanctot, M., and de Freitas, N. (2016) · 2016
Later among the works it cites.
Variational inference: A review for statisticians
Blei, D. M., Kucukelbir, A., and McAuliffe, J. D. (2017) · 2017
Closest in time.
Noisy network for exploration
Original
Fortunato, M., Azar, M. G., Piot, B., Menick, J., Osband, I., Graves, A., Mnih, V., Munos, R., Hassabis, D., Pietquin, I., Blundell, C., and Legg, S. (2017) · 2017
Closest in time.
Automatic differentiation variational inference
Kucukelbir, A., Tran, D., Ranganath, R., Gelman, A., and Blei, D. M. (2017) · 2017
Closest in time.
Deep probabilistic programming
Tran, D., Hoffman, M. D., Saurous, R. A., Brevdo, E., Murphy, K., and Blei, D. M. (2017) · 2017
Closest in time.