Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Later among the works it cites.
Bias in natural actor-critic algorithms
Philip Thomas · 2014
Later among the works it cites.
Taming the noise in reinforcement learning via soft updates
Original
Roy Fox, Ari Pakman, and Naftali Tishby · 2015
Later among the works it cites.
Policy gradient methods for off-policy control
Original
Lucas Lehnert and Doina Precup · 2015
Later among the works it cites.
End-to-end training of deep visuomotor policies
Original
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2015
Later among the works it cites.
Continuous control with deep reinforcement learning
Original
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Later among the works it cites.
Prioritized experience replay
Original
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2015
Later among the works it cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Later among the works it cites.
On-policy vs. off-policy updates for deep reinforcement learning
Matthew Hausknecht and Peter Stone · 2016
Closest in time.
Asynchronous methods for deep reinforcement learning
Original
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy P Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Closest in time.
Reward augmented maximum likelihood for neural structured prediction
Original
Mohammad Norouzi, Samy Bengio, Zhifeng Chen, Navdeep Jaitly, Mike Schuster, Yonghui Wu, and Dale Schuurmans · 2016
Closest in time.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Closest in time.
Deep reinforcement learning with double Q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Closest in time.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot, and Nando de Freitas · 2016
Closest in time.