Fetching the paper…
Reading the bibliography…
This paper develops the method of Continuous Pontryagin Differentiable Programming (Continuous PDP), which enables a robot to learn an objective function from a few sparsely demonstrated keyframes.
N. Das, S. Bechtle, T. Davchev, D. Jayaraman, A. Rai, and F. Meier, “Model-based inverse reinforcement learning from visual demonstrations,” in Conference on Robotic Learning , pp. 1930–1942
1942
Earlier work this paper cites.
L. S. Pontryagin, V. G. Boltyanskiy, R. V. Gamkrelidze, and E. F. Mishchenko, The Mathematical Theory of Optimal Processes . John Wiley & Sons, Inc., 1962
1962
Earlier work this paper cites.
R. Bellman, “Dynamic programming,” Science , vol. 153, no. 3731, pp. 34–37, 1966
1966
Earlier work this paper cites.
D. H. Jacobson and D. Q. Mayne, Differential dynamic programming . Elsevier Publishing Company, 1970, no. 24
1970
Earlier work this paper cites.
P. Moylan and B. Anderson, “Nonlinear regulator theory and an inverse optimal control problem,” IEEE Transactions on Automatic Control , vol. 18, no. 5, pp. 460–465, 1973
1973
Earlier work this paper cites.
A. V. Fiacco, “Sensitivity analysis for nonlinear programming using penalty methods,” Mathematical programming , vol. 10, no. 1, pp. 287–311, 1976
1976
Earlier work this paper cites.
H. Sakoe and S. Chiba, “Dynamic programming algorithm optimization for spoken word recognition,” IEEE Transactions on Acoustics, Speech, and Signal Processing , vol. 26, no. 1, pp. 43–49, 1978
1978
Earlier work this paper cites.
C. D. Kolstad and L. S. Lasdon, “Derivative evaluation and computational experience with large bilevel mathematical programs,” Journal of optimization theory and applications , vol. 65, no. 3, pp. 485–499, 1990
1990
Earlier work this paper cites.
D. A. Pomerleau, “Efficient training of artificial neural networks for autonomous navigation,” Neural computation , vol. 3, no. 1, pp. 88–97, 1991
1991
Earlier work this paper cites.
P. Hansen, B. Jaumard, and G. Savard, “New branch-and-bound rules for linear bilevel programming,” SIAM Journal on scientific and Statistical Computing , vol. 13, no. 5, pp. 1194–1217, 1992
1992
Earlier work this paper cites.
J. B. Kuipers, Quaternions and rotation sequences . Princeton University Press, 1999, vol. 66
1999
Earlier work this paper cites.
A. Y. Ng, S. J. Russell et al. , “Algorithms for inverse reinforcement learning.” in International Conference of Machine Learning , vol. 1, 2000, p. 2
2000
Earlier work this paper cites.
P. Abbeel and A. Y. Ng, “Apprenticeship learning via inverse reinforcement learning,” in International Conference on Machine Learning , 2004, pp. 1–8
2004
Earlier work this paper cites.
W. Li and E. Todorov, “Iterative linear quadratic regulator design for nonlinear biological movement systems,” in International Conference on Informatics in Control, Automation and Robotics , vol. 2, 2004, pp. 222–229
2004
Earlier work this paper cites.
N. D. Ratliff, J. A. Bagnell, and M. A. Zinkevich, “Maximum margin planning,” in International Conference on Machine Learning , 2006, pp. 729–736
2006
Earlier work this paper cites.
A. Wächter and L. T. Biegler, “On the implementation of an interior-point filter line-search algorithm for large-scale nonlinear programming,” Mathematical programming , vol. 106, no. 1, pp. 25–57, 2006
2006
Earlier work this paper cites.
S. Calinon, F. Guenter, and A. Billard, “On learning, representing, and generalizing a task in a humanoid robot,” IEEE Transactions on Systems, Man, and Cybernetics , vol. 37, no. 2, pp. 286–298, 2007
2007
Earlier work this paper cites.
B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey, “Maximum entropy inverse reinforcement learning.” in Association for the Advancement of Artificial Intelligence , vol. 8, 2008, pp. 1433–1438
2008
Earlier work this paper cites.
M. W. Spong and M. Vidyasagar, Robot dynamics and control . John Wiley & Sons, 2008
2008
Earlier work this paper cites.
M. J. Powell, “The bobyqa algorithm for bound constrained optimization without derivatives,” Cambridge NA Report, University of Cambridge, Cambridge , pp. 26–46, 2009
2009
Earlier work this paper cites.
K. Mombaur, A. Truong, and J.-P. Laumond, “From human to humanoid locomotion—an inverse optimal control approach,” Autonomous Robots , vol. 28, no. 3, pp. 369–383, 2010
2010
Cited alongside, same era.
T. Lee, M. Leok, and N. H. McClamroch, “Geometric tracking control of a quadrotor uav on se(3),” in IEEE Conference on Decision and Control , 2010, pp. 5420–5425
2010
Cited alongside, same era.
A. Keshavarz, Y. Wang, and S. Boyd, “Imputing a convex objective function,” in IEEE International Symposium on Intelligent Control . IEEE, 2011, pp. 613–619
2011
Cited alongside, same era.
A.-S. Puydupin-Jamin, M. Johnson, and T. Bretl, “A convex approach to inverse optimal control and its application to modeling human locomotion,” in International Conference on Robotics and Automation , 2012, pp. 531–536
2012
Cited alongside, same era.
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, M. Kudlur, J. Levenberg, R. Monga, S. Moore, D. G. Murray, B. Steiner, P. Tucker, V. Vasudevan, P. Warden, M. Wicke, Y. Yu, and X. Zheng, “TensorFlow: A system for Large-Scale machine learning,” in 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16) . USENIX Association, 2016, pp. 265–283
2016
Later among the works it cites.
P. Englert, N. A. Vien, and M. Toussaint, “Inverse kkt: Learning cost functions of manipulation tasks from demonstrations,” International Journal of Robotics Research , vol. 36, no. 13-14, pp. 1474–1488, 2017
2017
Later among the works it cites.
A. Sinha, P. Malo, and K. Deb, “A review on bilevel optimization: from classical to evolutionary approaches and applications,” IEEE Transactions on Evolutionary Computation , vol. 22, no. 2, pp. 276–295, 2017
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Hatz, J. P. Schloder, and H. G. Bock, “Estimating parameters in optimal control problems,” SIAM Journal on Scientific Computing , vol. 34, no. 3, pp. A1707–A1728, 2012
2012
Cited alongside, same era.
A. Vakanski, I. Mantegh, A. Irish, and F. Janabi-Sharifi, “Trajectory learning for robot programming by demonstration using hidden markov model and dynamic time warping,” IEEE Transactions on Systems, Man, and Cybernetics , vol. 42, no. 4, pp. 1039–1052, 2012
2012
Cited alongside, same era.
B. Akgun, M. Cakmak, K. Jiang, and A. L. Thomaz, “Keyframe-based learning from demonstration,” International Journal of Social Robotics , vol. 4, no. 4, pp. 343–355, 2012
2012
Cited alongside, same era.
B. Akgun, M. Cakmak, J. W. Yoo, and A. L. Thomaz, “Trajectories and keyframes for kinesthetic teaching: A human-robot interaction perspective,” in ACM/IEEE international conference on Human-Robot Interaction , 2012, pp. 391–398
2012
Cited alongside, same era.
S. G. Krantz and H. R. Parks, The implicit function theorem: history, theory, and applications . Springer Science & Business Media, 2012
2012
Cited alongside, same era.
F. L. Lewis, D. Vrabie, and V. L. Syrmos, Optimal control . John Wiley & Sons, 2012
2012
Cited alongside, same era.
P. Englert, A. Paraschos, M. P. Deisenroth, and J. Peters, “Probabilistic model-based imitation learning,” Adaptive Behavior , vol. 21, no. 5, pp. 388–403, 2013
2013
Cited alongside, same era.
K. Mombaur, A.-H. Olivier, and A. Crétual, “Forward and inverse optimal control of bipedal running,” in Modeling, simulation and optimization of bipedal walking . Springer, 2013, pp. 165–179
2013
Cited alongside, same era.
2017
Later among the works it cites.
C. Moro, G. Nejat, and A. Mihailidis, “Learning and personalizing socially assistive robot behaviors to aid with activities of daily living,” ACM Transactions on Human-Robot Interaction , vol. 7, no. 2, pp. 1–25, 2018
2018
Later among the works it cites.
R. Rahmatizadeh, P. Abolghasemi, L. Bölöni, and S. Levine, “Vision-based multi-task manipulation for inexpensive robots using end-to-end learning from demonstration,” in IEEE International Conference on Robotics and Automation , 2018, pp. 3758–3765
2018
Later among the works it cites.
F. Torabi, G. Warnell, and P. Stone, “Behavioral cloning from observation,” in International Joint Conference on Artificial Intelligence , 2018, pp. 4950–4957
2018
Later among the works it cites.
T. Osa, J. Pajarinen, G. Neumann, J. A. Bagnell, P. Abbeel, J. Peters et al. , “An algorithmic perspective on imitation learning,” Foundations and Trends in Robotics , vol. 7, no. 1-2, pp. 1–179, 2018
2018
Later among the works it cites.
B. Amos, I. Jimenez, J. Sacks, B. Boots, and J. Z. Kolter, “Differentiable mpc for end-to-end planning and control,” in Advances in Neural Information Processing Systems , 2018, pp. 8299––8310
2018
Later among the works it cites.
C. C. Aggarwal et al. , “Neural networks and deep learning,” Springer , vol. 10, pp. 978–3, 2018
2018
Later among the works it cites.
W. Jin, D. Kulić, J. F.-S. Lin, S. Mou, and S. Hirche, “Inverse optimal control for multiphase cost functions,” IEEE Transactions on Robotics , vol. 35, no. 6, pp. 1387–1398, 2019
2019
Later among the works it cites.
C.-Y. Chang, D.-A. Huang, Y. Sui, L. Fei-Fei, and J. C. Niebles, “D3tw: Discriminative differentiable dynamic time warping for weakly supervised action alignment and segmentation,” in IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 3546–3555
2019
Later among the works it cites.
J. A. E. Andersson, J. Gillis, G. Horn, J. B. Rawlings, and M. Diehl, “CasADi – A software framework for nonlinear optimization and optimal control,” Mathematical Programming Computation , vol. 11, no. 1, pp. 1–36, 2019
2019
Later among the works it cites.
H. Ravichandar, A. S. Polydoros, S. Chernova, and A. Billard, “Recent advances in robot learning from demonstration,” Annual Review of Control, Robotics, and Autonomous Systems , vol. 3, 2020
2020
Closest in time.
W. Jin, Z. Wang, Z. Yang, and S. Mou, “Pontryagin differentiable programming: An end-to-end learning and control framework,” in Advances in Neural Information Processing Systems , 2020
2020
Closest in time.
W. Jin, D. Kulić, S. Mou, and S. Hirche, “Inverse optimal control from incomplete trajectory observations,” International Journal of Robotics Research , vol. 40, no. 6-7, pp. 848–865, 2021
2021
Closest in time.
W. Jin and S. Mou, “Distributed inverse optimal control,” Automatica , vol. 129, p. 109658, 2021
2021
Closest in time.
W. Jin, S. Mou, and G. J. Pappas, “Safe pontryagin differentiable programming,” in Advances in Neural Information Processing Systems , 2021
2021
Closest in time.
K. Ji, J. Yang, and Y. Liang, “Bilevel optimization: Convergence analysis and enhanced design,” in International Conference on Machine Learning , 2021, pp. 4882–4892
2021
Closest in time.
Z. Liang, W. Jin, and S. Mou, “An iterative method for inverse optimal control,” in Asian Control Conference , 2022, pp. 959–964
2022
Closest in time.