Fetching the paper…
Reading the bibliography…
We focus on the problem of learning a single motor module that can flexibly express a range of behaviors for the control of high-dimensional physically simulated humanoids.
A second-order gradient method for determining optimal trajectories of non-linear discrete-time systems
David Mayne · 1966
Earlier work this paper cites.
Differential Dynamic Programming
David H. Jacobson and David Q. Mayne · 1970
Earlier work this paper cites.
The role and use of the stochastic linear-quadratic-gaussian problem in control system design
Michael Athans · 1971
Earlier work this paper cites.
Minimax differential dynamic programming: An application to robust biped walking
Jun Morimoto and Christopher G Atkeson · 2003
Earlier work this paper cites.
Learning movement primitives
Stefan Schaal, Jan Peters, Jun Nakanishi, and Auke Ijspeert · 2003
Earlier work this paper cites.
Unsupervised learning of sensory-motor primitives
Emanuel Todorov and Zoubin Ghahramani · 2003
Earlier work this paper cites.
The organization of behavioral repertoire in motor cortex
Michael Graziano · 2006
Earlier work this paper cites.
Combining modules for movement
E Bizzi, VCK Cheung, A d’Avella, P Saltiel, and Matthew Tresch · 2008
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol · 2008
Earlier work this paper cites.
LQR-trees: Feedback motion planning on sparse randomized trees
Russ Tedrake · 2009
Earlier work this paper cites.
Sampling-based contact-rich motion control
Libin Liu, KangKang Yin, Michiel van de Panne, Tianjia Shao, and Weiwei Xu · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Synthesis and stabilization of complex behaviors through online trajectory optimization
Yuval Tassa, Tom Erez, and Emanuel Todorov · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Probabilistic movement primitives
Alexandros Paraschos, Christian Daniel, Jan R Peters, and Gerhard Neumann · 2013
Cited alongside, same era.
Learning modular policies for robotics
Gerhard Neumann, Christian Daniel, Alexandros Paraschos, Andras Kupcsik, and Jan Peters · 2014
Cited alongside, same era.
Stochastic backpropagation and approximate inference in deep generative models
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra · 2014
Cited alongside, same era.
Control-limited differential dynamic programming
Yuval Tassa, Nicolas Mansard, and Emo Todorov · 2014
Cited alongside, same era.
Learning reduced-order feedback policies for motion skills
Kai Ding, Libin Liu, Michiel Van de Panne, and KangKang Yin · 2015
Cited alongside, same era.
Learning continuous control policies by stochastic value gradients
Nicolas Heess, Gregory Wayne, David Silver, Tim Lillicrap, Tom Erez, and Yuval Tassa · 2015
Sobolev training for neural networks
Wojciech M Czarnecki, Simon Osindero, Max Jaderberg, Grzegorz Swirszcz, and Razvan Pascanu · 2017
Later among the works it cites.
One-shot imitation learning
Yan Duan, Marcin Andrychowicz, Bradly Stadie, OpenAI Jonathan Ho, Jonas Schneider, Ilya Sutskever, Pieter Abbeel, and Wojciech Zaremba · 2017
Later among the works it cites.
Emergence of locomotion behaviours in rich environments
Nicolas Heess, Srinivasan Sriram, Jay Lemmon, Josh Merel, Greg Wayne, Yuval Tassa, Tom Erez, Ziyu Wang, Ali Eslami, Martin Riedmiller, et al · 2017
Later among the works it cites.
Dart: Noise injection for robust imitation learning
Michael Laskey, Jonathan Lee, Roy Fox, Anca Dragan, and Ken Goldberg · 2017
Later among the works it cites.
Learning human behaviors from motion capture by adversarial imitation
Josh Merel, Yuval Tassa, Sriram Srinivasan, Jay Lemmon, Ziyu Wang, Greg Wayne, and Nicolas Heess · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Interactive control of diverse complex characters with neural networks
Igor Mordatch, Kendall Lowrey, Galen Andrew, Zoran Popovic, and Emanuel V Todorov · 2015
Cited alongside, same era.
Actor-mimic: Deep multitask and transfer reinforcement learning
Emilio Parisotto, Jimmy Lei Ba, and Ruslan Salakhutdinov · 2015
Cited alongside, same era.
Andrei A Rusu, Sergio Gomez Colmenarejo, Caglar Gulcehre, Guillaume Desjardins, James Kirkpatrick, Razvan Pascanu, Volodymyr Mnih, Koray Kavukcuoglu, and Raia Hadsell · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Deeploco: Dynamic locomotion skills using hierarchical deep reinforcement learning
Xue Bin Peng, Glen Berseth, KangKang Yin, and Michiel Van De Panne · 2017
Later among the works it cites.
Distral: Robust multitask reinforcement learning
Yee Whye Teh, Victor Bapst, Wojciech M Czarnecki, John Quan, James Kirkpatrick, Raia Hadsell, Nicolas Heess, and Razvan Pascanu · 2017
Later among the works it cites.
Emergent complexity via multi-agent competition
Trapit Bansal, Jakub Pachocki, Szymon Sidor, Ilya Sutskever, and Igor Mordatch · 2018
Closest in time.
Physics-based motion capture imitation with deep reinforcement learning
Nuttapong Chentanez, Matthias Müller, Miles Macklin, Viktor Makoviychuk, and Stefan Jeschke · 2018
Closest in time.
Tommaso Furlanello, Zachary C Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar · 2018
Closest in time.
Learning basketball dribbling skills using trajectory optimization and deep reinforcement learning
Libin Liu and Jessica Hodgins · 2018
Closest in time.
Hierarchical visuomotor control of humanoids
Josh Merel, Arun Ahuja, Vu Pham, Saran Tunyasuvunakool, Siqi Liu, Dhruva Tirumala, Nicolas Heess, and Greg Wayne · 2018
Closest in time.
Deepmimic: Example-guided deep reinforcement learning of physics-based character skills
Xue Bin Peng, Pieter Abbeel, Sergey Levine, and Michiel van de Panne · 2018
Closest in time.
Knowledge transfer with jacobian matching
Suraj Srinivas and François Fleuret · 2018
Closest in time.
Robust imitation of diverse behaviors
Ziyu Wang, Josh S Merel, Scott E Reed, Nando de Freitas, Gregory Wayne, and Nicolas Heess · 2018
Closest in time.