Bayesian reinforcement learning: A survey
Ghavamzadeh, M., Mannor, S., Pineau, J., Tamar, A., et al. (2015) · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Later among the works it cites.
Exploratory gradient boosting for reinforcement learning in complex domains
Abel, D., Agarwal, A., Diaz, F., Krishnamurthy, A., and Schapire, R. E. (2016) · 2016
Later among the works it cites.
Unifying count-based exploration and intrinsic motivation
Bellemare, M., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R. (2016) · 2016
Later among the works it cites.
Openai gym
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Later among the works it cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z. (2016) · 2016
Later among the works it cites.
Contextual decision processes with low bellman rank are pac-learnable
Jiang, N., Krishnamurthy, A., Agarwal, A., Langford, J., and Schapire, R. E. (2016) · 2016
Later among the works it cites.
End-to-end training of deep visuomotor policies
Levine et al., S. (2016) · 2016
Later among the works it cites.
Deep exploration via bootstrapped dqn
Osband, I., Blundell, C., Pritzel, A., and Van Roy, B. (2016) · 2016
Later among the works it cites.
Safe, multi-agent, reinforcement learning for autonomous driving
Shalev-Shwartz, S., Shammah, S., and Shashua, A. (2016) · 2016
Later among the works it cites.
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D. (2016) · 2016
Later among the works it cites.
Linear thompson sampling revisited
Abeille, M. and Lazaric, A. (2017) · 2017
Later among the works it cites.
Noisy networks for exploration
Fortunato, M., Azar, M. G., Piot, B., Menick, J., Osband, I., Graves, A., Mnih, V., Munos, R., Hassabis, D., Pietquin, O., et al. (2017) · 2017
Later among the works it cites.
Deep reinforcement learning that matters
Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., and Meger, D. (2017) · 2017
Later among the works it cites.
Shallow updates for deep reinforcement learning
Levine, N., Zahavy, T., Mankowitz, D. J., Tamar, A., and Mannor, S. (2017) · 2017
Later among the works it cites.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Original
Machado, M. C., Bellemare, M. G., Talvitie, E., Veness, J., Hausknecht, M., and Bowling, M. (2017) · 2017
Later among the works it cites.
Count-based exploration with neural density models
Ostrovski, G., Bellemare, M. G., Oord, A. v. d., and Munos, R. (2017) · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al. (2017) · 2017
Later among the works it cites.
Stochastic activation pruning for robust adversarial defense
Original
Dhillon, G. S., Azizzadenesheli, K., Lipton, Z. C., Bernstein, J., Kossaifi, J., Khanna, A., and Anandkumar, A. (2018) · 2018
Closest in time.