Fetching the paper…
Reading the bibliography…
We identify a fundamental problem in policy gradient-based methods in continuous control.
On the convergence of policy iteration in stationary dynamic programming
Martin L Puterman and Shelby L Brumelle · 1979
Earlier work this paper cites.
New stochastic approximation type procedures
Boris T Polyak · 1990
Earlier work this paper cites.
Robust estimation of a location parameter
Peter J Huber · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 1994
Earlier work this paper cites.
No free lunch theorems for optimization
David H Wolpert, William G Macready, et al · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Actor-critic algorithms
Vijay R Konda and John N Tsitsiklis · 2000
Earlier work this paper cites.
Quantile regression: An introduction
Roger Koenker and Kevin Hallock · 2001
Earlier work this paper cites.
Foundations of modern probability
Olav Kallenberg · 2006
Earlier work this paper cites.
Quantile autoregression
Roger Koenker and Zhijie Xiao · 2006
Earlier work this paper cites.
Stochastic approximation: a dynamical systems viewpoint , volume 48
Vivek S Borkar · 2009
Earlier work this paper cites.
An online actor–critic algorithm with function approximation for constrained markov decision processes
Shalabh Bhatnagar and K Lakshmanan · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Bayesian policy gradient and actor-critic algorithms
Mohammad Ghavamzadeh, Yaakov Engel, and Michal Valko · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Evolution strategies as a scalable alternative to reinforcement learning
Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever · 2017
Later among the works it cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
Learning dexterous in-hand manipulation
Marcin Andrychowicz, Bowen Baker, Maciek Chociej, Rafal Jozefowicz, Bob McGrew, Jakub Pachocki, Arthur Petron, Matthias Plappert, Glenn Powell, Alex Ray, et al · 2018
Later among the works it cites.
Addressing function approximation error in actor-critic methods
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Cited alongside, same era.
Improved techniques for training gans
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen · 2016
Cited alongside, same era.
Wasserstein generative adversarial networks
Martin Arjovsky, Soumith Chintala, and Léon Bottou · 2017
Cited alongside, same era.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Cited alongside, same era.
Improving stochastic policy gradients in continuous control with deep reinforcement learning using the beta distribution
Po-Wei Chou, Daniel Maturana, and Sebastian Scherer · 2017
Cited alongside, same era.
Risk-constrained reinforcement learning with percentile risk criteria
Yinlam Chow, Mohammad Ghavamzadeh, Lucas Janson, and Marco Pavone · 2017
Cited alongside, same era.
Distributional reinforcement learning with quantile regression
Will Dabney, Mark Rowland, Marc G Bellemare, and Rémi Munos · 2017
Cited alongside, same era.
Scott Fujimoto, Herke van Hoof, and David Meger · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Autoregressive quantile networks for generative modeling
Georg Ostrovski, Will Dabney, and Rémi Munos · 2018
Later among the works it cites.
Deepmimic: Example-guided deep reinforcement learning of physics-based character skills
Xue Bin Peng, Pieter Abbeel, Sergey Levine, and Michiel van de Panne · 2018
Later among the works it cites.
Nonlinear distributional gradient temporal-difference learning
Chao Qu, Shie Mannor, and Huan Xu · 2018
Later among the works it cites.
Learning by playing solving sparse reward tasks from scratch
Martin Riedmiller, Roland Hafner, Thomas Lampe, Michael Neunert, Jonas Degrave, Tom Wiele, Vlad Mnih, Nicolas Heess, and Jost Tobias Springenberg · 2018
Later among the works it cites.
An analysis of categorical distributional reinforcement learning
Mark Rowland, Marc G Bellemare, Will Dabney, Rémi Munos, and Yee Whye Teh · 2018
Later among the works it cites.
Autoregressive policies for continuous control deep reinforcement learning
Dmytro Korenkevych, A Rupam Mahmood, Gautham Vasan, and James Bergstra · 2019
Closest in time.
Discretizing continuous action space for on-policy optimization
Yunhao Tang and Shipra Agrawal · 2019
Closest in time.