The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G.; Naddaf, Y.; Veness, J.; and Bowling, M. 2013 · 2013
Cited alongside, same era.
Generative adversarial nets
Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; and Bengio, Y. 2014 · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A. A.; Veness, J.; Bellemare, M. G.; Graves, A.; Riedmiller, M.; Fidjeland, A. K.; Ostrovski, G.; et al. 2015 · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J.; Levine, S.; Abbeel, P.; Jordan, M.; and Moritz, P. 2015 · 2015
Cited alongside, same era.
Concrete problems in AI safety
Original
Amodei, D.; Olah, C.; Steinhardt, J.; Christiano, P.; Schulman, J.; and Mané, D. 2016 · 2016
Cited alongside, same era.
Guided cost learning: Deep inverse optimal control via policy optimization
Finn, C.; Levine, S.; and Abbeel, P. 2016 · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Ho, J.; and Ermon, S. 2016 · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Levine, S.; Finn, C.; Darrell, T.; and Abbeel, P. 2016 · 2016
Cited alongside, same era.
Unsupervised representation learning with deep convolutional generative adversarial networks
Radford, A.; Metz, L.; and Chintala, S. 2016 · 2016
Cited alongside, same era.
Deep reinforcement learning from human preferences
Christiano, P. F.; Leike, J.; Brown, T.; Martic, M.; Legg, S.; and Amodei, D. 2017 · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Original
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Cited alongside, same era.
Learning Robust Rewards with Adverserial Inverse Reinforcement Learning
Fu, J.; Luo, K.; and Levine, S. 2018 · 2018
Cited alongside, same era.