Taming the noise in reinforcement learning via soft updates
Original
Roy Fox, Ari Pakman, and Naftali Tishby · 2015
Later among the works it cites.
Learning continuous control policies by stochastic value gradients
Nicolas Heess, Gregory Wayne, David Silver, Timothy Lillicrap, Tom Erez, and Yuval Tassa · 2015
Later among the works it cites.
Continuous control with deep reinforcement learning
Original
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Later among the works it cites.
Universal value function approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Later among the works it cites.
Reward augmented maximum likelihood for neural structured prediction
Mohammad Norouzi, Samy Bengio, Navdeep Jaitly, Mike Schuster, Yonghui Wu, Dale Schuurmans, et al · 2016
Later among the works it cites.
Artificial intelligence: a modern approach
Stuart J Russell and Peter Norvig · 2016
Later among the works it cites.
Hindsight experience replay
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba · 2017
Later among the works it cites.
Bridging the gap between value and policy based reinforcement learning
Ofir Nachum, Mohammad Norouzi, Kelvin Xu, and Dale Schuurmans · 2017
Later among the works it cites.
Robust adversarial reinforcement learning
Lerrel Pinto, James Davidson, Rahul Sukthankar, and Abhinav Gupta · 2017
Later among the works it cites.
Maximum a posteriori policy optimisation
Original
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Remi Munos, Nicolas Heess, and Martin Riedmiller · 2018
Later among the works it cites.
Understanding the impact of entropy in policy learning
Original
Zafarali Ahmed, Nicolas Le Roux, Mohammad Norouzi, and Dale Schuurmans · 2018
Later among the works it cites.
Path consistency learning in tsallis entropy regularized mdps
Yinlam Chow, Ofir Nachum, and Mohammad Ghavamzadeh · 2018
Later among the works it cites.
Composable deep reinforcement learning for robotic manipulation
Tuomas Haarnoja, Vitchyr Pong, Aurick Zhou, Murtaza Dalal, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Original
Sergey Levine · 2018
Later among the works it cites.
Variational bayesian reinforcement learning with regret bounds
Original
Brendan O’Donoghue · 2018
Later among the works it cites.
Unsupervised control through non-parametric discriminative rewards
Original
David Warde-Farley, Tom Van de Wiele, Tejas Kulkarni, Catalin Ionescu, Steven Hansen, and Volodymyr Mnih · 2018
Later among the works it cites.
Adversarial examples are not bugs, they are features
Original
Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, and Aleksander Madry · 2019
Closest in time.
Skew-fit: State-covering self-supervised reinforcement learning
Original
Vitchyr H Pong, Murtaza Dalal, Steven Lin, Ashvin Nair, Shikhar Bahl, and Sergey Levine · 2019
Closest in time.
End-to-end robotic reinforcement learning without reward engineering
Original
Avi Singh, Larry Yang, Kristian Hartikainen, Chelsea Finn, and Sergey Levine · 2019
Closest in time.