Fetching the paper…
Reading the bibliography…
In environments with continuous state and action spaces, state-of-the-art actor-critic reinforcement learning algorithms can solve very complex problems, yet can also fail in environments that seem trivial, but the reason for such failures is still poorly understood.
Towards Characterizing Divergence in Deep Q-Learning
Joshua Achiam, Ethan Knight, and Pieter Abbeel · 1903
Earlier work this paper cites.
Learning with Delayed Rewards
Christopher J. C. H. Watkins · 1989
Earlier work this paper cites.
Reinforcement learning with high-dimensional, continuous actions
L. C. Baird and A. H. Klopf · 1993
Earlier work this paper cites.
Generalization in reinforcement learning: Safely approximating the value function
Justin A. Boyan and Andrew W. Moore · 1995
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
John N. Tsitsiklis and Benjamin Van Roy · 1997
Earlier work this paper cites.
Autonomous shaping: Knowledge transfer in reinforcement learning
George Konidaris and Andrew Barto · 2006
Earlier work this paper cites.
Reinforcement learning in continuous action spaces
Hado Van Hasselt and Marco A. Wiering · 2007
Earlier work this paper cites.
Parametric value function approximation: A unified view
M. Geist and O. Pietquin · 2011
Earlier work this paper cites.
Playing Atari with Deep Reinforcement Learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Cited alongside, same era.
Deterministic Policy Gradient Algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
Qt-Opt: Scalable deep reinforcement learning for vision-based robotic manipulation
Dmitry Kalashnikov, Alex Irpan, Peter Pastor, Julian Ibarz, Alexander Herzog, Eric Jang, Deirdre Quillen, Ethan Holly, Mrinal Kalakrishnan, Vincent Vanhoucke, et al · 2018
Later among the works it cites.
Learning by Playing - Solving Sparse Reward Tasks from Scratch
Martin Riedmiller, Roland Hafner, Thomas Lampe, Michael Neunert, Jonas Degrave, Tom Van de Wiele, Volodymyr Mnih, Nicolas Heess, and Jost Tobias Springenberg · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Later among the works it cites.
Deep Reinforcement Learning and the Deadly Triad
Hado van Hasselt, Yotam Doron, Florian Strub, Matteo Hessel, Nicolas Sonnerat, and Joseph Modayil · 2018
Later among the works it cites.
Understanding the impact of entropy on policy optimization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki, Tom Schaul, Joel Z Leibo, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Noisy Networks for Exploration
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Ian Osband, Alex Graves, Vlad Mnih, Remi Munos, Demis Hassabis, Olivier Pietquin, Charles Blundell, and Shane Legg · 2017
Cited alongside, same era.
Parameter space noise for exploration
Matthias Plappert, Rein Houthooft, Prafulla Dhariwal, Szymon Sidor, Richard Y. Chen, Xi Chen, Tamim Asfour, Pieter Abbeel, and Marcin Andrychowicz · 2017
Cited alongside, same era.
GEP-PG: Decoupling Exploration and Exploitation in Deep Reinforcement Learning Algorithms
Cédric Colas, Olivier Sigaud, and Pierre-Yves Oudeyer · 2018
Cited alongside, same era.
Off-Policy Deep Reinforcement Learning without Exploration
Scott Fujimoto, David Meger, and Doina Precup
Cited in the paper.
Addressing Function Approximation Error in Actor-Critic Methods
Scott Fujimoto, Herke van Hoof, and David Meger
Cited in the paper.
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine
Cited in the paper.
Soft actor-critic algorithms and applications
Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, George Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta, Pieter Abbeel, et al
Cited in the paper.
Zafarali Ahmed, Nicolas Le Roux, Mohammad Norouzi, and Dale Schuurmans · 2019
Closest in time.
Q-learning for continuous actions with cross-entropy guided policies
Riley Simmons-Edler, Ben Eisner, Eric Mitchell, Sebastian Seung, and Daniel Lee · 2019
Closest in time.
Exploiting the sign of the advantage function to learn deterministic policies in continuous domains
Matthieu Zimmer and Paul Weng · 2019
Closest in time.