Fetching the paper…
Reading the bibliography…
This work shows that policies with simple linear and RBF parameterizations can be trained to solve a variety of continuous control tasks, including the OpenAI gym benchmarks.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Leemon Baird · 1995
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-ichi Amari · 1998
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2001
Earlier work this paper cites.
Handbook of learning and approximate dynamic programming
Jennie Si · 2004
Earlier work this paper cites.
Machine learning of motor skills for robotics
Jan Peters · 2007
Earlier work this paper cites.
Random Features for Large-Scale Kernel Machines
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
Approximate dynamic programming
Dimitri P Bertsekas · 2008
Earlier work this paper cites.
Infinite-horizon model predictive control for periodic tasks with contacts
Tom Erez, Yuval Tassa, and Emanuel Todorov · 2011
Earlier work this paper cites.
Synthesis and stabilization of complex behaviors through online trajectory optimization
Yuval Tassa, Tom Erez, and Emanuel Todorov · 2012
Earlier work this paper cites.
Discovery of complex behaviors through contact-invariant optimization
Igor Mordatch, Emanuel Todorov, and Zoran Popovic · 2012
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
Trajectory Optimization for Full-Body Movements with Complex Contacts
Mazen Al Borno, Martin de Lasa, and Aaron Hertzmann · 2013
Cited alongside, same era.
An integrated system for real-time model predictive control of humanoid robots
Tom Erez, Kendall Lowrey, Yuval Tassa, Vikash Kumar, Svetoslav Kolev, and Emanuel Todorov · 2013
Cited alongside, same era.
A tutorial on linear function approximators for dynamic programming and reinforcement learning
Alborz Geramifard, Thomas J Walsh, Stefanie Tellex, Girish Chowdhary, Nicholas Roy, and Jonathan P How · 2013
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih et al · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Philipp Moritz, Michael Jordan, and Pieter Abbeel · 2015
Optimal control with learned local models: Application to dexterous manipulation
Vikash Kumar, Emanuel Todorov, and Sergey Levine · 2016
Later among the works it cites.
Learning dexterous manipulation policies from experience and imitation
Vikash Kumar, Abhishek Gupta, Emanuel Todorov, and Sergey Levine · 2016
Later among the works it cites.
Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours
Lerrel Pinto and Abhinav Gupta · 2016
Later among the works it cites.
Openai gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Later among the works it cites.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Ensemble-CIO: Full-body dynamic motion planning that transfers to physical humanoids
Igor Mordatch, Kendall Lowrey, and Emanuel Todorov · 2015
Cited alongside, same era.
Learning continuous control policies by stochastic value gradients
Nicolas Heess, Gregory Wayne, David Silver, Tim Lillicrap, Tom Erez, and Yuval Tassa · 2015
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver et al · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Cited alongside, same era.
Fereshteh Sadeghi and Sergey Levine · 2016
Later among the works it cites.
Benchmarking deep reinforcement learning for continuous control
Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel · 2016
Later among the works it cites.
Q-Prop: Sample-Efficient Policy Gradient with An Off-Policy Critic
Shixiang Gu, Timothy Lillicrap, Zoubin Ghahramani, Richard E. Turner, and Sergey Levine · 2017
Closest in time.
EPOpt: Learning Robust Neural Network Policies Using Model Ensembles
Aravind Rajeswaran, Sarvjeet Ghotra, Balaraman Ravindran, and Sergey Levine · 2017
Closest in time.
Domain randomization for transferring deep neural networks from simulation to the real world
Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel · 2017
Closest in time.