Fetching the paper…
Reading the bibliography…
Real-world robots are becoming increasingly complex and commonly act in poorly understood environments where it is extremely challenging to model or learn their true dynamics.
H. J. Kushner, “A new method of locating the maximum point of an arbitrary multipeak curve in the presence of noise,” Journal of Basic Engineering , vol. 86, p. 97, 1964
1964
Earlier work this paper cites.
J. Močkus, “On bayesian methods for seeking the extremum,” in Optimization Techniques IFIP Technical Conference , 1975
1975
Earlier work this paper cites.
M. Grimble, “Implicit and explicit LQG self-tuning controllers,” Automatica , vol. 20, no. 5, pp. 661–669, 1984
1984
Earlier work this paper cites.
D. Clarke, P. Kanjilal, and C. Mohtadi, “A generalized LQG approach to self-tuning control part i. aspects of design,” International Journal of Control , vol. 41, no. 6, pp. 1509–1523, 1985
1985
Earlier work this paper cites.
D. J. Bender and A. J. Laub, “The linear-quadratic optimal regulator for descriptor systems: discrete-time case,” Automatica , 1987
1987
Earlier work this paper cites.
H. Hjalmarsson, M. Gevers, and F. De Bruyne, “For model-based control design, closed-loop identification gives better performance,” Automatica , vol. 32, no. 12, 1996
1996
Earlier work this paper cites.
C. G. Atkeson, “Nonparametric model-based reinforcement learning,” in Advances in neural information processing systems , 1998, pp. 1008–1014
1998
Earlier work this paper cites.
R. Murray-Smith and D. Sbarbaro, “Nonlinear adaptive control using nonparametric Gaussian process prior models,” IFAC Proceedings Volumes , vol. 35, no. 1, pp. 325–330, 2002
2002
Earlier work this paper cites.
M. Gevers, “Identification for control: From the early achievements to the revival of experiment design,” European journal of control , vol. 11, no. 4-5, 2005
2005
Earlier work this paper cites.
J. Lofberg, “YALMIP: A toolbox for modeling and optimization in MATLAB,” in International Symposium on Computer Aided Control Systems Design , 2005, pp. 284–289
2005
Earlier work this paper cites.
P. Abbeel, M. Quigley, and A. Y. Ng, “Using inaccurate models in reinforcement learning,” in International conference on Machine learning, 2006 , pp. 1–8
2006
Cited alongside, same era.
C. E. Rasmussen and C. K. I. Williams, Gaussian Processes for Machine Learning . The MIT Press, 2006
2006
Cited alongside, same era.
M. A. Osborne, R. Garnett, and S. J. Roberts, “Gaussian processes for global optimization,” in Learning and Intelligent Optimization (LION3) , 2009, pp. 1–15
2009
Cited alongside, same era.
D. Nguyen-Tuong and J. Peters, “Model learning for robot control: a survey,” Cognitive Processing , vol. 12, no. 4, pp. 319–340, 2011
2011
Cited alongside, same era.
J. W. Roberts, I. R. Manchester, and R. Tedrake, “Feedback controller parameterizations for reinforcement learning,” in Symposium on Adaptive Dynamic Programming And Reinforcement Learning (ADPRL) , 2011, pp. 310–317
R. Martinez-Cantin, “BayesOpt: a Bayesian optimization library for nonlinear optimization, experimental design and bandits.” Journal of Machine Learning Research , vol. 15, no. 1, pp. 3735–3739, 2014
2014
Later among the works it cites.
M. P. Deisenroth, D. Fox, and C. E. Rasmussen, “Gaussian processes for data-efficient learning in robotics and control,” Transactions on Pattern Analysis and Machine Intelligence (PAMI) , 2015
2015
Later among the works it cites.
R. Calandra, A. Seyfarth, J. Peters, and M. P. Deisenroth, “Bayesian optimization for learning gaits under uncertainty,” Annals of Mathematics and Artificial Intelligence , vol. 76, no. 1, pp. 5–23, 2015
2015
Later among the works it cites.
B. Landry, “Planning and control for quadrotor flight through cluttered environments,” Master’s thesis, MIT, 2015
2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2011
Cited alongside, same era.
S. Sastry and M. Bodson, Adaptive control: stability, convergence and robustness . Courier Corporation, 2011
2011
Cited alongside, same era.
N. Abas, A. Legowo, and R. Akmeliawati, “Parameter identification of an autonomous quadrotor,” in International Conference On Mechatronics , 2011, pp. 1–8
2011
Cited alongside, same era.
J. Joseph, A. Geramifard, J. W. Roberts, J. P. How, and N. Roy, “Reinforcement learning with misspecified model classes,” in International Conference on Robotics and Automation , 2013, pp. 939–946
2013
Cited alongside, same era.
K. J. Åström and B. Wittenmark, Adaptive control . Courier Corporation, 2013
2013
Cited alongside, same era.
S. Trimpe, A. Millane, S. Doessegger, and R. D’Andrea, “A self-tuning LQR approach demonstrated on an inverted pendulum,” IFAC Proceedings Volumes , vol. 47, no. 3, pp. 11 281–11 287, 2014
2014
Cited alongside, same era.
Y. Sui, A. Gotovos, J. Burdick, and A. Krause, “Safe exploration for optimization with Gaussian processes,” in International Conference on Machine Learning , 2015, pp. 997–1005
2015
Later among the works it cites.
B. Shahriari, K. Swersky, Z. Wang, R. P. Adams, and N. de Freitas, “Taking the human out of the loop: A review of Bayesian optimization,” Proceedings of the IEEE , vol. 104, no. 1, pp. 148–175, 2016
2016
Later among the works it cites.
A. Marco, P. Hennig, J. Bohg, S. Schaal, and S. Trimpe, “Automatic LQR tuning based on Gaussian process global optimization,” in International Conference on Robotics and Automation , 2016
2016
Later among the works it cites.
S. Bansal, A. K. Akametalu, F. J. Jiang, F. Laine, and C. J. Tomlin, “Learning quadrotor dynamics using neural network for flight control,” in Conference on Decision and Control , 2016, pp. 4653–4660
2016
Later among the works it cites.
F. Berkenkamp, A. P. Schoellig, and A. Krause, “Safe controller optimization for quadrotors with Gaussian processes,” in International Conference on Robotics and Automation, 2016 , pp. 491–496
2016
Later among the works it cites.