Fetching the paper…
Reading the bibliography…
Reinforcement Learning (RL) has demonstrated a huge potential in learning optimal policies without any prior knowledge of the process to be controlled.
H. Bock and K. Plitt, “A multiple shooting algorithm for direct solution of optimal control problems,” in Proceedings 9th IFAC World Congress Budapest . Pergamon Press, 1984, pp. 242–247
1984
Earlier work this paper cites.
C. Watkins, “Learning from delayed rewards,” Ph.D. dissertation, King’s College, Cambridge, 1989
1989
Earlier work this paper cites.
F. Y. Wang and I. T. Cameron, “Control studies on a model evaporation process — constrained state driving with conventional and higher relative degree systems,” Journal of Process Control , vol. 4, pp. 59–75, 1994
1994
Earlier work this paper cites.
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” in Proceedings of the 12th International Conference on Neural Information Processing Systems , ser. NIPS’99. Cambridge, MA, USA: MIT Press, 1999, pp. 1057–1063
1999
Earlier work this paper cites.
P. Scokaert and J. Rawlings, “Feasibility Issues in Linear Model Predictive Control,” AIChE Journal , vol. 45, no. 8, pp. 1649–1659, 1999
1999
Earlier work this paper cites.
C. Büskens and H. Maurer, Online Optimization of Large Scale Systems . Berlin, Heidelberg: Springer Berlin Heidelberg, 2001, ch. Sensitivity Analysis and Real-Time Optimization of Parametric Nonlinear Programming Problems, pp. 3–16
2001
Earlier work this paper cites.
C. Sonntag, O. Stursberg, and S. Engell, “Dynamic Optimization of an Industrial Evaporator using Graph Search with Embedded Nonlinear Programming,” in Proc. 2nd IFAC Conf. on Analysis and Design of Hybrid Systems (ADHS) , 2006, pp. 211–216
2006
Earlier work this paper cites.
P. Abbeel, A. Coates, M. Quigley, and A. Y. Ng, “An application of reinforcement learning to aerobatic helicopter flight,” in In Advances in Neural Information Processing Systems 19 . MIT Press, 2007, p. 2007
2007
Earlier work this paper cites.
J. Rawlings and D. Mayne, Model Predictive Control: Theory and Design . Nob Hill, 2009
2009
Earlier work this paper cites.
F. L. Lewis and D. Vrabie, “Reinforcement learning and adaptive dynamic programming for feedback control,” IEEE Circuits and Systems Magazine , vol. 9, no. 3, pp. 32–50, 2009
2009
Cited alongside, same era.
L. Grüne and J. Pannek, Nonlinear Model Predictive Control . London: Springer, 2011
2011
Cited alongside, same era.
S. Wang, W. Chaovalitwongse, and R. Babuska, “Machine learning algorithms in bipedal robot control,” IEEE Transactions on Systems, Man, and Cybernetics Part C , vol. 42, no. 5, pp. 728–743, Sep. 2012
2012
Cited alongside, same era.
F. L. Lewis, D. Vrabie, and K. G. Vamvoudakis, “Reinforcement learning and feedback control: Using natural decision methods to design optimal adaptive controllers,” IEEE Control Systems , vol. 32, no. 6, pp. 76–105, 2012
2012
Cited alongside, same era.
G. Theocharous, P. S. Thomas, and M. Ghavamzadeh, “Personalized Ad Recommendation Systems for Life-Time Value Optimization with Guarantees,” in IJCAI , 2015, pp. 1806–1812
2015
Later among the works it cites.
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis, “Mastering the game of go with deep neural networks and tree search,” Nature , vol. 529, pp. 484–503, 2016
2016
Later among the works it cites.
C. J. Ostafew, A. P. Schoellig, and T. D. Barfoot, “Robust Constrained Learning-based NMPC enabling reliable mobile robot path tracking,” The International Journal of Robotics Research , vol. 35, no. 13, pp. 1547–1563, 2016
2016
Later among the works it cites.
M. Zanon, S. Gros, and M. Diehl, “A Tracking MPC Formulation that is Locally Equivalent to Economic MPC,” Journal of Process Control , 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2013
Cited alongside, same era.
J. P. Jens Kober, J. Andrew Bagnell, “Reinforcement learning in robotics: A survey,” The International Journal of Robotics Research , vol. 32, 2013
2013
Cited alongside, same era.
R. Amrit, J. B. Rawlings, and L. T. Biegler, “Optimizing process economics online using model predictive control,” Computers & Chemical Engineering , vol. 58, pp. 334 – 343, 2013
2013
Cited alongside, same era.
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller, “Deterministic policy gradient algorithms,” in Proceedings of the 31st International Conference on International Conference on Machine Learning - Volume 32 , ser. ICML’14, 2014, pp. I–387–I–395
2014
Cited alongside, same era.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski et al. , “Human-level control through deep reinforcement learning,” Nature , vol. 518, no. 7540, p. 529, 2015
2015
Cited alongside, same era.
2016
Later among the works it cites.
F. Borrelli, A. Bemporad, and M. Morari, Predictive control for linear and hybrid systems . Cambridge University Press, 2017
2017
Later among the works it cites.
F. Berkenkamp, M. Turchetta, A. Schoellig, and A. Krause, “Safe Model-based Reinforcement Learning with Stability Guarantees,” in Advances in Neural Information Processing Systems 30 , I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, Eds. Curran Associates, Inc., 2017, pp. 908–918
2017
Later among the works it cites.
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction. Second Edition . MIT press Cambridge, 2018. [Online]. Available: http://incompleteideas.net/book/the-book-2nd.html
2018
Later among the works it cites.
T. Koller, F. Berkenkamp, M. Turchetta, and A. Krause, “Learning-based Model Predictive Control for Safe Exploration and Reinforcement Learning,” 2018, published on Arxiv
2018
Later among the works it cites.
S. Gros and M. Zanon, “Data-Driven Economic NMPC using Reinforcement Learning,” IEEE Transactions on Automatic Control , 2018, (under revision). [Online]. Available: https://mariozanon.wordpress.com/rlfornmpc
2018
Later among the works it cites.