Fetching the paper…
Reading the bibliography…
The process of learning a manipulation task depends strongly on the action space used for exploration: posed in the incorrect action space, solving a task with reinforcement learning can be drastically inefficient.
M. T. Mason, “Compliance and force control for computer controlled manipulators,” Transactions on Systems, Man, and Cybernetics , vol. 11, no. 6, pp. 418–432, 6 1981
1981
Earlier work this paper cites.
N. Hogan, “Impedance control: An approach to manipulation,” Journal of dynamic systems, measurement, and control , vol. 107, p. 17, 1985
1985
Earlier work this paper cites.
O. Khatib, “A unified approach for motion and force control of robot manipulators: The operational space formulation,” IEEE Journal on Robotics and Automation , vol. 3, no. 1, pp. 43–53, 1987
1987
Earlier work this paper cites.
H. Bruyninckx and J. De Schutter, “Specification of force-controlled actions in the "task frame formalism"-a synthesis,” Transactions on Robotics and Automation , vol. 12, no. 4, pp. 581–589, 8 1996
1996
Earlier work this paper cites.
R. S. Sutton, D. Precup, and S. Singh, “Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning,” Artificial intelligence , vol. 112, no. 1-2, pp. 181–211, 1999
1999
Earlier work this paper cites.
A. J. Ijspeert, J. Nakanishi, and S. Schaal, “Movement imitation with nonlinear dynamical systems in humanoid robots,” in ICRA , vol. 2, 5 2002, pp. 1398–1403 vol.2
2002
Earlier work this paper cites.
M. Stolle and D. Precup, “Learning options in reinforcement learning,” in International Symposium on abstraction, reformulation, and approximation . Springer, 2002, pp. 212–223
2002
Earlier work this paper cites.
I. Menache, S. Mannor, and N. Shimkin, “Q-cut—dynamic discovery of sub-goals in reinforcement learning,” in European Conference on Machine Learning . Springer, 2002, pp. 295–306
2002
Earlier work this paper cites.
T. Kröger, B. Finkemeyer, U. Thomas, and F. M. Wahl, “Compliant motion programming: The task frame formalism revisited,” Mechatronics & Robotics, Aachen, Germany , 2004
2004
Earlier work this paper cites.
J. Buchli, F. Stulp, E. Theodorou, and S. Schaal, “Learning variable impedance control,” IJRR , vol. 30, no. 7, pp. 820–833, 2011
2011
Earlier work this paper cites.
G. Konidaris, S. Kuindersma, R. Grupen, and A. Barto, “Autonomous skill acquisition on a mobile manipulator,” in Twenty-Fifth AAAI Conference on Artificial Intelligence , 2011
2011
Earlier work this paper cites.
P. Baldi, “Autoencoders, unsupervised learning, and deep architectures,” in Proceedings of ICML workshop on unsupervised and transfer learning , 2012, pp. 37–49
2012
Earlier work this paper cites.
2013
Cited alongside, same era.
S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,” JMLR , vol. 17, no. 1, 2016
2016
Cited alongside, same era.
M. Khansari, E. Klingbeil, and O. Khatib, “Adaptive human-inspired compliant contact primitives to perform surface–surface contact under uncertainty,” IJRR , vol. 35, no. 13, pp. 1651–1675, 2016
2016
Cited alongside, same era.
T. D. Kulkarni, K. Narasimhan, A. Saeedi, and J. Tenenbaum, “Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation,” in Advances in neural information processing systems , 2016, pp. 3675–3683
2016
Cited alongside, same era.
2018
Later among the works it cites.
R. Martín-Martín, M. Lee, R. Gardner, S. Savarese, J. Bohg, and A. Garg, “Variable impedance control in end-effector space. an action space for reinforcement learning in contact rich tasks,” in Proceedings of the International Conference of Intelligent Robots and Systems (IROS) , 2019
2019
Later among the works it cites.
P. Varin, L. Grossman, and S. Kuindersma, “A comparison of action spaces for learning manipulation tasks,” in Proceedings of the International Conference of Intelligent Robots and Systems (IROS) , 2019
2019
Later among the works it cites.
2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
B. Siciliano and O. Khatib, Springer handbook of robotics . Springer, 2016
2016
Cited alongside, same era.
A. Pervez, Y. Mao, and D. Lee, “Learning deep movement primitives using convolutional neural networks,” in 2017 IEEE-RAS 17th International Conference on Humanoid Robotics (Humanoids) . IEEE, 2017, pp. 191–197
2017
Cited alongside, same era.
P.-L. Bacon, J. Harb, and D. Precup, “The option-critic architecture,” in Thirty-First AAAI Conference on Artificial Intelligence , 2017
2017
Cited alongside, same era.
S. Krishnan*, A. Garg*, S. Patil, C. Lea, G. Hager, P. Abbeel, and K. Goldberg (* equal contribution), “Transition state clustering: Unsupervised surgical trajectory segmentation for robot learning,” IJRR , vol. 36, no. 13-14, pp. 1595–1618, 2017
2017
Cited alongside, same era.
C. Finn, P. Abbeel, and S. Levine, “Model-agnostic meta-learning for fast adaptation of deep networks,” in Proceedings of the 34th International Conference on Machine Learning-Volume 70 , 2017, pp. 1126–1135
2017
Cited alongside, same era.
I. Higgins, L. Matthey, A. Pal, C. Burgess, X. Glorot, M. Botvinick, S. Mohamed, and A. Lerchner, “beta-vae: Learning basic visual concepts with a constrained variational framework.” Iclr , vol. 2, no. 5, p. 6, 2017
2017
Cited alongside, same era.
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction . MIT press, 2018
2018
Cited alongside, same era.
Later among the works it cites.
K. Fang, Y. Zhu, A. Garg, S. Savarese, and L. Fei-Fei, “Dynamics learning with cascaded variational inference for multi-step manipulation,” in Conference on Robot Learning (CoRL) , oct 2019
2019
Later among the works it cites.
W. Whitney, R. Agarwal, K. Cho, and A. Gupta, “Dynamics-aware embeddings,” in International Conference on Learning Representations , 2019
2019
Later among the works it cites.
Y. Chandak, G. Theocharous, J. Kostas, S. Jordan, and P. Thomas, “Learning action representations for reinforcement learning,” in International Conference on Machine Learning , 2019, pp. 941–950
2019
Later among the works it cites.
2019
Later among the works it cites.
M. Botvinick, S. Ritter, J. X. Wang, Z. Kurth-Nelson, C. Blundell, and D. Hassabis, “Reinforcement learning, fast and slow,” Trends in cognitive sciences , vol. 23, no. 5, pp. 408–422, 2019
2019
Later among the works it cites.
J. Gao, Y. Zhou, and T. Asfour, “Learning compliance adaptation in contact-rich manipulation,” in Proceedings of the International Conference on Robotics and Automotion (ICRA) , 2020
2020
Later among the works it cites.
E. van der Pol, T. Kipf, F. A. Oliehoek, and M. Welling, “Plannable approximations to mdp homomorphisms: Equivariance under actions,” in Proceedings of the 19th International Conference on Autonomous Agents and MultiAgent Systems , 2020, pp. 1431–1439
2020
Later among the works it cites.