Fetching the paper…
Reading the bibliography…
The naive application of Reinforcement Learning algorithms to continuous control problems -- such as locomotion and manipulation -- often results in policies which rely on high-amplitude, high-frequency control signals, known colloquially as bang-bang control.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Introduction to Reinforcement Learning
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Constrained Markov Decision Processes
E. Altman · 1999
Earlier work this paper cites.
A geometric approach to multi-criterion reinforcement learning
Shie Mannor and Nahum Shimkin · 2004
Earlier work this paper cites.
Relative entropy policy search
Jan Peters and Katharina Mülling · 2010
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Guided policy search
Sergey Levine and Vladlen Koltun · 2013
Earlier work this paper cites.
A survey of multi-objective sequential decision-making
Diederik M. Roijers, Peter Vamplew, Shimon Whiteson, and Richard Dazeley · 2013
Earlier work this paper cites.
Multi-objective optimization
Kalyanmoy Deb · 2014
Cited alongside, same era.
The Penn Jerboa: A platform for exploring parallel composition of templates
Avik De and Daniel E. Koditschek · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Cited alongside, same era.
Design principles for a family of direct-drive legged robots
G. Kenneally, A. De, and D. E. Koditschek · 2016
Cited alongside, same era.
Data-efficient deep reinforcement learning for dexterous manipulation
Ivaylo Popov, Nicolas Heess, Timothy P. Lillicrap, Roland Hafner, Gabriel Barth-Maron, Matej Vecerik, Thomas Lampe, Yuval Tassa, Tom Erez, and Martin A. Riedmiller · 2017
Later among the works it cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel · 2017
Later among the works it cites.
Maximum a Posteriori Policy Optimisation
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Remi Munos, Nicolas Heess, and Martin Riedmiller · 2018
Later among the works it cites.
Safe exploration in continuous action spaces
Gal Dalal, Krishnamurthy Dvijotham, Matej Vecerik, Todd Hester, Cosmin Paduraru, and Yuval Tassa · 2018
Later among the works it cites.
Sim-to-real: Learning agile locomotion for quadruped robots
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hossam Mossalam, Yannis M. Assael, Diederik M. Roijers, and Shimon Whiteson · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
Rémi Munos, Tom Stepleton, Anna Harutyunyan, and Marc G. Bellemare · 2016
Cited alongside, same era.
Constrained Policy Optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Cited alongside, same era.
Jie Tan, Tingnan Zhang, Erwin Coumans, Atil Iscen, Yunfei Bai, Danijar Hafner, Steven Bohez, and Vincent Vanhoucke · 2018
Later among the works it cites.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, Timothy P. Lillicrap, and Martin A. Riedmiller · 2018
Later among the works it cites.
Reward constrained policy optimization
Chen Tessler, Daniel J. Mankowitz, and Shie Mannor · 2018
Later among the works it cites.