Policy gradient methods for reinforcement learning with function approximation
Sutton, R., McAllester, D. A., Singh, S., and Mansour, Y · 1999
Earlier work this paper cites.
Invariant causal prediction for block mdps
Original
Zhang, A., Lyle, C., Sodhani, S., Filos, A., Kwiatkowska, M., Pineau, J., Gal, Y., and Precup, D · 2003
Earlier work this paper cites.
Learning invariant representations for reinforcement learning without reconstruction
Original
Zhang, A., McAllister, R., Calandra, R., Gal, Y., and Levine, S · 2006
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G. E., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Original
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Original
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M. I., and Moritz, P · 2015
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Original
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Original
Wang, Z., Schaul, T., Hessel, M., Hasselt, H. V., Lanctot, M., and Freitas, N. D · 2016
Earlier work this paper cites.
Unsupervised learning of disentangled representations from video
Denton, E. L. and Birodkar, V · 2017
Earlier work this paper cites.
The kinetics human action video dataset
Original
Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, A., Suleyman, M., and Zisserman, A · 2017
Earlier work this paper cites.
Towards generalization and simplicity in continuous control
Original
Rajeswaran, A., Lowrey, K., Todorov, E., and Kakade, S. M · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Original
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., Chen, Y., Lillicrap, T., Hui, F., Sifre, L., van den Driessche, G., Graepel, T., and Hassabis, D · 2017
Earlier work this paper cites.
A study on overfitting in deep reinforcement learning
Original
Zhang, C., Vinyals, O., Munos, R., and Bengio, S · 2017
Earlier work this paper cites.
Distributed distributional deterministic policy gradients
Original
Barth-Maron, G., Hoffman, M. W., Budden, D., Dabney, W., Horgan, D., Dhruva, T., Muldal, A., Heess, N., and Lillicrap, T · 2018
Earlier work this paper cites.
Quantifying generalization in reinforcement learning
Original
Cobbe, K., Klimov, O., Hesse, C., Kim, T., and Schulman, J · 2018
Earlier work this paper cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Original
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., et al · 2018
Earlier work this paper cites.