Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Later among the works it cites.
Policy gradient for coherent risk measures
Aviv Tamar, Yinlam Chow, Mohammad Ghavamzadeh, and Shie Mannor · 2015
Later among the works it cites.
OpenAI Gym
Original
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Later among the works it cites.
Bayesian Policy Gradient and Actor-Critic Algorithms
Mohammad Ghavamzadeh, Yaakov Engel, and Michal Valko · 2016
Later among the works it cites.
Continuous Control with Deep Reinforcement Learning
Original
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Later among the works it cites.
Reinforcement Learning in Robust Markov Decision Processes
Shiau Hong Lim, Huan Xu, and Shie Mannor · 2016
Later among the works it cites.
Robust MDPs with k-Rectangular Uncertainty
Shie Mannor, Ofir Mebel, and Huan Xu · 2016
Later among the works it cites.
Distributionally Robust Counterpart in Markov Decision Processes
Pengqian Yu and Huan Xu · 2016
Later among the works it cites.
Deep Robust Kalman Filter
Original
Shirli Di-Castro Shashua and Shie Mannor · 2017
Later among the works it cites.
Reinforcement learning under Model Mismatch
Aurko Roy, Huan Xu, and Sebastian Pokutta · 2017
Later among the works it cites.
Learning Robust Options
Daniel J Mankowitz, Timothy A Mann, Shie Mannor, Doina Precup, and Pierre-Luc Bacon · 2018
Closest in time.