Fetching the paper…
Reading the bibliography…
We present a new method of learning control policies that successfully operate under unknown dynamic models.
Evolutionary Robotics: The Biology, Intelligence, and Technology
S. Nolfi and D. Floreano · 2000
Earlier work this paper cites.
ε \varepsilon -MDPs: Learning in varying environments
István Szita, Bálint Takács, and András Lörincz · 2002
Earlier work this paper cites.
Back-to-Reality: Crossing the Reality Gap in Evolutionary Robotics
Juan Cristobal Zagal, Javier Ruiz-del-Solar, and Paul Vallejos · 2004
Earlier work this paper cites.
Exploration and Apprenticeship Learning in Reinforcement Learning
Pieter Abbeel and Andrew Y. Ng · 2005
Earlier work this paper cites.
Nonlinear System Identification Using Coevolution of Models and Tests
Josh C. Bongard and Hod Lipson · 2005
Earlier work this paper cites.
Using Inaccurate Models in Reinforcement Learning
Pieter Abbeel, Morgan Quigley, and Andrew Y. Ng · 2006
Earlier work this paper cites.
System identification without lennart ljung: what would have been different?
Michel Gevers · 2006
Earlier work this paper cites.
Crossing the reality gap in evolutionary robotics by promoting transferable controllers
Sylvain Koos, Jean-Baptiste Mouret, and Stéphane Doncieux · 2010
Earlier work this paper cites.
PILCO: A model-based and data-efficient approach to policy search
Marc Deisenroth and Carl E Rasmussen · 2011
Earlier work this paper cites.
Bruno Da Silva, George Konidaris, and Andrew Barto · 2012
Earlier work this paper cites.
Agnostic system identification for model-based reinforcement learning
Stephane Ross and J. Andrew Bagnell · 2012
Cited alongside, same era.
Learning compact parameterized skills with a single regression
Freek Stulp, Gennaro Raiola, Antoine Hoarau, Serena Ivaldi, and Olivier Sigaud · 2013
Cited alongside, same era.
Deterministic Policy Gradient Algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin A. Riedmiller · 2014
Cited alongside, same era.
Reducing Hardware Experiments for Model Learning and Policy Optimization
Sehoon Ha and Katsu Yamane · 2015
Cited alongside, same era.
Memory-based control with recurrent neural networks
Nicolas Heess, Jonathan J Hunt, Timothy P Lillicrap, and David Silver · 2015
Cited alongside, same era.
Deep learning helicopter dynamics models
Ali Punjani and Pieter Abbeel · 2015
Later among the works it cites.
Transfer from Simulation to Real World through Learning Deep Inverse Dynamics Model
Paul Christiano, Zain Shah, Igor Mordatch, Jonas Schneider, Trevor Blackwell, Joshua Tobin, Pieter Abbeel, and Wojciech Zaremba · 2016
Later among the works it cites.
Q-Prop: Sample-Efficient Policy Gradient with An Off-Policy Critic
Shixiang Gu, Timothy Lillicrap, Zoubin Ghahramani, Richard E Turner, and Sergey Levine · 2016
Later among the works it cites.
3D Simulation for Robot Arm Control with Deep Q-Learning
Stephen James and Edward Johns · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sergey Levine, Nolan Wagener, and Pieter Abbeel · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Ensemble-CIO: Full-body dynamic motion planning that transfers to physical humanoids
Igor Mordatch, Kendall Lowrey, and Emanuel Todorov · 2015
Cited alongside, same era.
URL http://dartsim.github.io/
DART: Dynamic Animation and Robotics Toolkit”
Cited in the paper.
URL https://github.com/sehoonha/pydart2
Pydart2
Cited in the paper.
End-to-End Training of Deep Visuomotor Policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel
Cited in the paper.
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy P Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Later among the works it cites.
Terrain-Adaptive Locomotion Skills Using Deep Reinforcement Learning
Xue Bin Peng, Glen Berseth, and Michiel van de Panne · 2016
Later among the works it cites.
Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours
Lerrel Pinto and Abhinav Gupta · 2016
Later among the works it cites.
EPOpt: Learning Robust Neural Network Policies Using Model Ensembles
Aravind Rajeswaran, Sarvjeet Ghotra, Sergey Levine, and Balaraman Ravindran · 2016
Later among the works it cites.
Sim-to-real robot learning from pixels with progressive nets
Andrei A Rusu, Matej Vecerik, Thomas Rothörl, Nicolas Heess, Razvan Pascanu, and Raia Hadsell · 2016
Later among the works it cites.