Fetching the paper…
Reading the bibliography…
Dexterous multi-fingered hands are extremely versatile and provide a generic way to perform a multitude of tasks in human-centric environments.
ALVINN: an autonomous land vehicle in a neural network
D. Pomerleau · 1988
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Natural gradient works efficiently in learning
S. Amari · 1998
Earlier work this paper cites.
A natural policy gradient
S. Kakade · 2001
Earlier work this paper cites.
Movement imitation with nonlinear dynamical systems in humanoid robots
A. J. Ijspeert, J. Nakanishi, and S. Schaal · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
S. Kakade and J. Langford · 2002
Earlier work this paper cites.
Covariant policy search
J. A. Bagnell and J. G. Schneider · 2003
Earlier work this paper cites.
Computational approaches to motor learning by imitation
S. Schaal, A. Ijspeert, and A. Billard · 2003
Earlier work this paper cites.
Machine learning of motor skills for robotics
J. Peters · 2007
Earlier work this paper cites.
Natural actor-critic
J. Peters and S. Schaal · 2007
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
J. Peters and S. Schaal · 2008
Earlier work this paper cites.
Learning motor primitives for robotics
J. Kober and J. Peters · 2009
Earlier work this paper cites.
A generalized path integral control approach to reinforcement learning
E. Theodorou, J. Buchli, and S. Schaal · 2010
Earlier work this paper cites.
Reinforcement learning of motor skills in high dimensions: A path integral approach
E. Theodorou, J. Buchli, and S. Schaal · 2010
Earlier work this paper cites.
Policy search for motor primitives in robotics
J. Kober and J. Peters · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. J. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
Integrating reinforcement learning with human demonstrations of varying ability
M. E. Taylor, H. B. Suay, and S. Chernova · 2011
Earlier work this paper cites.
Generalization of human grasping for multi-fingered robot hands
H. B. Amor, O. Kroemer, U. Hillenbrand, G. Neumann, and J. Peters · 2012
Earlier work this paper cites.
Contact-invariant optimization for hand manipulation
I. Mordatch, Z. Popović, and E. Todorov · 2012
Cited alongside, same era.
MuJoCo: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Cited alongside, same era.
Fast, strong and compliant pneumatic actuation for dexterous tendon-driven hands
V. Kumar, Z. Xu, and E. Todorov · 2013
Cited alongside, same era.
A low-cost and modular, 20-dof anthropomorphic robotic hand: design, actuation and modeling
Z. Xu, V. Kumar, and E. Todorov · 2013
Cited alongside, same era.
Real-time behaviour synthesis for dynamic hand-manipulation
V. Kumar, Y. Tassa, T. Erez, and E. Todorov · 2014
Cited alongside, same era.
A direct method for trajectory optimization of rigid bodies through contact
M. Posa, C. Cantu, and R. Tedrake · 2014
Cited alongside, same era.
Learning Dexterous Manipulation Policies from Experience and Imitation
V. Kumar, A. Gupta, E. Todorov, and S. Levine · 2016
Later among the works it cites.
Optimal control with learned local models: Application to dexterous manipulation
V. Kumar, E. Todorov, and S. Levine · 2016
Later among the works it cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Later among the works it cites.
Exploration from demonstration for interactive reinforcement learning
K. Subramanian, C. L. I. Jr., and A. L. Thomaz · 2016
Later among the works it cites.
Hindsight experience replay
M. Andrychowicz, D. Crow, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, P. Abbeel, and W. Zaremba · 2017
Closest in time.
Deep predictive policy training using reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforcement learning from demonstration through shaping
T. Brys, A. Harutyunyan, H. B. Suay, S. Chernova, M. E. Taylor, and A. Nowé · 2015
Cited alongside, same era.
Simulation tools for model-based robotics: Comparison of bullet, havok, mujoco, ode and physx
T. Erez, Y. Tassa, and E. Todorov · 2015
Cited alongside, same era.
Mujoco haptix: A virtual reality system for hand manipulation
V. Kumar and E. Todorov · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Cited alongside, same era.
Ensemble-CIO: Full-body dynamic motion planning that transfers to physical humanoids
I. Mordatch, K. Lowrey, and E.Todorov · 2015
Cited alongside, same era.
Trust region policy optimization
J. Schulman, S. Levine, P. Moritz, M. Jordan, and P. Abbeel · 2015
Cited alongside, same era.
A. Ghadirzadeh, A. Maki, D. Kragic, and M. Björkman · 2017
Closest in time.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates
S. Gu, E. Holly, T. P. Lillicrap, and S. Levine · 2017
Closest in time.
Emergence of locomotion behaviours in rich environments
N. Heess, D. TB, S. Sriram, J. Lemmon, J. Merel, G. Wayne, Y. Tassa, T. Erez, Z. Wang, S. M. A. Eslami, M. A. Riedmiller, and D. Silver · 2017
Closest in time.
Deep reinforcement learning that matters
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger · 2017
Closest in time.
Learning from demonstrations for real world reinforcement learning
T. Hester, M. Vecerik, O. Pietquin, M. Lanctot, T. Schaul, B. Piot, A. Sendonaris, G. Dulac-Arnold, I. Osband, J. Agapiou, J. Z. Leibo, and A. Gruslys · 2017
Closest in time.
Overcoming exploration in reinforcement learning with demonstrations
A. Nair, B. McGrew, M. Andrychowicz, W. Zaremba, and P. Abbeel · 2017
Closest in time.
EPOpt: Learning Robust Neural Network Policies Using Model Ensembles
A. Rajeswaran, S. Ghotra, B. Ravindran, and S. Levine · 2017
Closest in time.
Towards Generalization and Simplicity in Continuous Control
A. Rajeswaran, K. Lowrey, E. Todorov, and S. Kakade · 2017
Closest in time.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Closest in time.
Deeply aggrevated: Differentiable imitation learning for sequential prediction
W. Sun, A. Venkatraman, G. J. Gordon, B. Boots, and J. A. Bagnell · 2017
Closest in time.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
M. Vecerik, T. Hester, J. Scholz, F. Wang, O. Pietquin, B. Piot, N. Heess, T. Rothörl, T. Lampe, and M. A. Riedmiller · 2017
Closest in time.
Reinforcement learning from imperfect demonstrations
Y. Gao, H. Xu, J. Lin, F. Yu, S. Levine, and T. Darrell · 2018
Closest in time.
Deepmimic: Example-guided deep reinforcement learning of physics-based character skills
X. B. Peng, P. Abbeel, S. Levine, and M. van de Panne · 2018
Closest in time.