Fetching the paper…
Reading the bibliography…
In this paper, we present a robotic model-based reinforcement learning method that combines ideas from model identification and model predictive control.
D. H. Jacobson and D. Q. Mayne, Differential dynamic programming , ser. Modern analytic and computational methods in science and mathematics. American Elsevier Pub. Co., 1970. [Online]. Available: http://books.google.com/books?id=tA-oAAAAIAAJ
1970
Earlier work this paper cites.
B. Armstrong, “On finding exciting trajectories for identification experiments involving systems with nonlinear dynamics,” Int. Journal of Robotics Research , vol. 8, no. 6, pp. 28–48, 1989
1989
Earlier work this paper cites.
M. Gautier and W. Khalil, “Exciting trajectories for the identification of base inertial parameters of robots,” Int. Journal of Robotics Research , vol. 11, no. 4, pp. 362–375, 1992
1992
Earlier work this paper cites.
R. S. Sutton, A. G. Barto, and R. J. Williams, “Reinforcement learning is direct adaptive optimal control,” IEEE Control Systems , vol. 12, no. 2, pp. 19–22, 1992
1992
Earlier work this paper cites.
K. J. Astrom and B. Wittenmark, Adaptive Control , 2nd ed. Boston, MA, USA: Addison-Wesley Longman Publishing Co., Inc., 1994
1994
Earlier work this paper cites.
J. Swevers, C. Ganseman, D. B. Tukel, J. De Schutter, and H. Van Brussel, “Optimal robot excitation and identification,” IEEE Trans. on Robotics and Automation (TRA) , vol. 13, no. 5, pp. 730–740, 1997
1997
Earlier work this paper cites.
L. Ljung, System identification . Springer, 1998
1998
Earlier work this paper cites.
P. Abbeel, A. Coates, M. Quigley, and A. Y. Ng, “An application of reinforcement learning to aerobatic helicopter flight,” in Advances in Neural Information Processing Systems , 2006
2006
Earlier work this paper cites.
P. L. B. Ambuj Tewari, “Optimistic linear programming gives logarithmic regret for irreducible MDPs,” in Proc. of Neural Information Processing Systems Conference (NIPS) , 2007
2007
Earlier work this paper cites.
K. P. Murphy, “Conjugate bayesian analysis of the gaussian distribution,” def , vol. 1, p. 16, 2007
2007
Earlier work this paper cites.
J. Hollerbach, W. Khalil, and M. Gautier, “Model identification,” in Springer Handbook of Robotics . Springer, 2008, pp. 321–344
2008
Earlier work this paper cites.
J.-Y. Audibert, R. Munos, and C. Szepesvári, “Exploration-exploitation tradeoff using variance estimates in multi-armed bandits,” Theoretical Computer Science , vol. 410, no. 19, pp. 1876–1902, 2009
2009
Cited alongside, same era.
D. Nguyen-Tuong and J. Peters, “Using model knowledge for learning inverse dynamics,” in 2014 IEEE International Conference on Robotics and Automation (ICRA) , 2010, pp. 2677–2682
2010
Cited alongside, same era.
2010
Cited alongside, same era.
Y. Abbasi-Yadkori, C. Szepesvári, S. Kakade, and U. V. Luxburg, “Regret bounds for the adaptive control of linear quadratic systems,” in Proc. of the 24th Annual Conference on Learning Theory , 2011
2011
Cited alongside, same era.
J. Boedecker, J. Springenberg, J. Wulfing, and M. Riedmiller, “Approximate real-time optimal control based on sparse gaussian process models,” in Adaptive Dynamic Programming and Reinforcement Learning (ADPRL) , 2014
2014
Later among the works it cites.
J. Boedecker, J. T. Springenberg, J. Wulfing, and M. Riedmiller, “Approximate real-time optimal control based on sparse gaussian process models,” in Adaptive Dynamic Programming and Reinforcement Learning (ADPRL), 2014 IEEE Symposium on . IEEE, 2014, pp. 1–8
2014
Later among the works it cites.
M. Cutler, T. Walsh, and J. How, “Reinforcement learning with multi-fidelity simulators,” in 2014 IEEE International Conference on Robotics and Automation (ICRA) , 2014, pp. 3888–3895
2014
Later among the works it cites.
C. J. Ostafew, A. P. Schoellig, and T. D. Barfoot, “Learning-based nonlinear model predictive control to improve vision-based mobile robot path-tracking in challenging outdoor environments.” in Proc. IEEE Int. Conf. Robotics and Automation (ICRA) , 2014
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. P. Deisenroth and C. E. Rasmussen, “PILCO: A model-based and data-efficient approach to policy search,” in Proceedings of the 28th International Conference on Machine Learning , L. Getoor and T. Scheffer, Eds. Bellevue, Washington, USA: Omnipress, 2011, pp. 465–472
2011
Cited alongside, same era.
M. Araya, O. Buffet, and V. Thomas, “Near-optimal BRL using optimistic local transitions,” in Proc. Int. Conf. on Machine Learning (ICML) , ser. ICML ’12, 2012, pp. 97–104
2012
Cited alongside, same era.
Y. Tassa, T. Erez, and E. Todorov, “Synthesis and stabilization of complex behaviors through online trajectory optimization,” in IEEE/RSJ International Conference on Intelligent Robots and Systems , 2012
2012
Cited alongside, same era.
E. F. Camacho and C. B. Alba, Model predictive control , 2013
2013
Cited alongside, same era.
M. Gautier, S. Briot, and G. Venture, “Identification of consistent standard dynamic parameters of industrial robots,” in IEEE/ASME Int. Conf. on Advanced Intelligent Mechatronics (AIM) , 2013, pp. 1429–1435
2013
Cited alongside, same era.
J. Kober, J. A. Bagnell, and J. Peters, “Reinforcement learning in robotics: A survey,” International Journal of Robotic Research , vol. 32, no. 11, pp. 1238–1274, 2013
2013
Cited alongside, same era.
S. Kuindersma, R. Grupen, and A. Barto, “Variational bayesian optimization for runtime risk-sensitive control,” in Robotics: Science and Systems (RSS) , 2013
2013
Cited alongside, same era.
2014
Later among the works it cites.
C. D. Sousa and R. Cortesao, “Physical feasibility of robot base inertial parameter identification: A linear matrix inequality approach,” Int. Journal of Robotics Research , vol. 33, no. 6, pp. 931–944, 2014
2014
Later among the works it cites.
Y. Tassa, N. Mansard, and E. Todorov, “Control-limited differential dynamic programming,” in International Conference on Robotics and Automation (ICRA) , 2014
2014
Later among the works it cites.
C. Wang, Y. Zhao, C.-Y. Lin, and M. Tomizuka, “Fast planning of well conditioned trajectories for model learning,” in IEEE/RSJ Int. Conf. on Intelligent Robots and Systems (IROS) , 2014, pp. 1460–1465
2014
Later among the works it cites.
M. Cutler and J. P. How, “Efficient reinforcement learning for robots using informative simulated priors,” in IEEE International Conference on Robotics and Automation (ICRA) , 2015
2015
Closest in time.
T. Moldovan, S. Levine, M. Jordan, and P. Abbeel, “Optimism-driven exploration for nonlinear systems,” in International Conference on Robotics and Automation (ICRA) , 2015
2015
Closest in time.
W. Rackl, R. Lampariello, and G. Hirzinger, “Robot excitation trajectories for dynamic parameter estimation using optimized b-splines,” in Proc. IEEE Int. Conf. Robotics and Automation (ICRA) , 2012, pp. 2042–2047
2047
Closest in time.