Stochastic backpropagation and approximate inference in deep generative models
Original
Rezende, D. J., Mohamed, S., and Wierstra, D · 2014
Cited alongside, same era.
Taming the noise in reinforcement learning via soft updates
Original
Fox, R., Pakman, A., and Tishby, N · 2015
Cited alongside, same era.
Made: Masked autoencoder for distribution estimation
Germain, M., Gregor, K., Murray, I., and Larochelle, H · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Original
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Cited alongside, same era.
Increasing the action gap: New operators for reinforcement learning
Bellemare, M. G., Ostrovski, G., Guez, A., Thomas, P. S., and Munos, R · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Van Hasselt, H., Guez, A., and Silver, D · 2016
Cited alongside, same era.
Input convex neural networks
Amos, B., Xu, L., and Kolter, J. Z · 2017
Cited alongside, same era.
Ucb exploration via q-ensembles
Original
Chen, R. Y., Sidor, S., Abbeel, P., and Schulman, J · 2017
Cited alongside, same era.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates
Gu, S., Holly, E., Lillicrap, T., and Levine, S · 2017
Cited alongside, same era.
Sequence tutor: Conservative fine-tuning of sequence generation models with kl-control
Jaques, N., Gu, S., Bahdanau, D., Hernández-Lobato, J. M., Turner, R. E., and Eck, D · 2017
Cited alongside, same era.