Fetching the paper…
Reading the bibliography…
In practice, the parameters of control policies are often tuned manually.
R. Calandra, A. Seyfarth, J. Peters, and M. P. Deisenroth, “An experimental comparison of Bayesian optimization for bipedal locomotion,” in IEEE International Conference on Robotics and Automation , 2014, pp. 1951–1958
1958
Earlier work this paper cites.
J. Mockus, Bayesian Approach to Global Optimization , ser. Mathematics and Its Applications, M. Hazewinkel, Ed. Springer, 1989, vol. 37
1989
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction . MIT press, 1998
1998
Earlier work this paper cites.
M. C. Kennedy and A. O’Hagan, “Predicting the output from a complex computer code when fast approximations are available,” Biometrika , vol. 87, no. 1, pp. 1–13, 2000
2000
Earlier work this paper cites.
J. Peters and S. Schaal, “Policy gradient methods for robotics,” in IEEE International Conference on Intelligent Robots and Systems , 2006, pp. 2219–2225
2006
Earlier work this paper cites.
P. Abbeel, M. Quigley, and A. Y. Ng, “Using Inaccurate Models in Reinforcement Learning,” in ACM International Conference on Machine Learning , 2006, pp. 1–8
2006
Earlier work this paper cites.
C. E. Rasmussen and C. K. Williams, Gaussian Processes for Machine Learning . MIT Press, 2006
2006
Earlier work this paper cites.
D. J. Lizotte, T. Wang, M. H. Bowling, and D. Schuurmans, “Automatic gait optimization with Gaussian process regression.” in International Joint Conference on Artificial Intelligence , vol. 7, 2007, pp. 944–949
2007
Earlier work this paper cites.
A. I. J. Forrester, A. Sóbester, and A. J. Keane, “Multi-fidelity optimization via surrogate modelling,” Royal Society of London A: Mathematical, Physical and Engineering Sciences , vol. 463, no. 2088, pp. 3251–3269, 2007
2007
Earlier work this paper cites.
B. D. O. Anderson and J. B. Moore, Optimal Control: Linear Quadratic Methods . Mineola, New York: Dover Publications, 2007
2007
Earlier work this paper cites.
J. Villemonteix, E. Vazquez, and E. Walter, “An informational approach to the global optimization of expensive-to-evaluate functions,” Journal of Global Optimization , vol. 44, no. 4, pp. 509–540, 2008
2008
Cited alongside, same era.
M. E. Taylor and P. Stone, “Transfer learning for reinforcement learning domains: A survey,” Journal of Machine Learning Research , vol. 10, pp. 1633–1685, 2009
2009
Cited alongside, same era.
P. Frazier, W. Powell, and S. Dayanik, “The knowledge-gradient policy for correlated normal beliefs,” INFORMS Journal on Computing , vol. 21, no. 4, pp. 599–613, 2009
2009
Cited alongside, same era.
N. Srinivas, A. Krause, S. M. Kakade, and M. Seeger, “Gaussian process optimization in the bandit setting: No regret and experimental design,” in International Conference on Machine Learning , 2010, pp. 1015–1022
2010
Cited alongside, same era.
J. Kober, J. A. Bagnell, and J. Peters, “Reinforcement learning in robotics: A survey,” The International Journal of Robotics Research , vol. 32, no. 11, pp. 1238–1274, 2013
2013
Later among the works it cites.
S. Trimpe, A. Millane, S. Doessegger, and R. D’Andrea, “A self-tuning LQR approach demonstrated on an inverted pendulum,” in 19th IFAC World Congress , 2014, pp. 11 281–11 287
2014
Later among the works it cites.
M. Cutler and J. P. How, “Efficient reinforcement learning for robots using informative simulated priors,” in IEEE International Conference on Robotics and Automation , 2015, pp. 2605–2612
2015
Later among the works it cites.
M. Cutler, T. J. Walsh, and J. P. How, “Real-world reinforcement learning via multifidelity simulators,” IEEE Transactions on Robotics , vol. 31, no. 3, pp. 655–671, 2015
2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Tesch, J. Schneider, and H. Choset, “Using response surfaces and expected improvement to optimize snake robot gait parameters,” in IEEE International Conference on Intelligent Robots and Systems , 2011, pp. 1069–1074
2011
Cited alongside, same era.
A. Krause and C. S. Ong, “Contextual Gaussian process bandit optimization,” in Neural Information Processing Systems , 2011, pp. 2447–2455
2011
Cited alongside, same era.
J. Roberts, I. Manchester, and R. Tedrake, “Feedback controller parameterizations for reinforcement learning,” in IEEE Symposium on Adaptive Dynamic Programming And Reinforcement Learning , 2011, pp. 310–317
2011
Cited alongside, same era.
P. Hennig and C. J. Schuler, “Entropy search for information-efficient global optimization,” Journal of Machine Learning Research , vol. 13, no. 1, pp. 1809–1837, 2012
2012
Cited alongside, same era.
K. Swersky, J. Snoek, and R. P. Adams, “Multi-Task Bayesian Optimization,” in Advances in Neural Information Processing Systems , 2013, pp. 2004–2012
2012
Cited alongside, same era.
Quanser, “Self-erecting single inverted pendulum – Instructor manual”,” Tech. Rep. 516, rev. 4.1
Cited in the paper.
A. Marco, P. Hennig, J. Bohg, S. Schaal, and S. Trimpe, “Automatic LQR tuning based on Gaussian process global optimization,” in IEEE International Conference on Robotics and Automation , 2016, pp. 270–277
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
H. Abdelrahman, F. Berkenkamp, J. Poland, and A. Krause, “Bayesian optimization for maximum power point tracking in photovoltaic power plants,” in European Control Conference , 2016, pp. 2078–2083
2083
Closest in time.