Fetching the paper…
Reading the bibliography…
Reinforcement Learning (RL) has recently impressed the world with stunning results in various applications.
D. Bertsekas and I. Rhodes, “Recursive state estimation for a set-membership description of uncertainty,” IEEE Transactions on Automatic Control , vol. 16, pp. 117–128, 1971
1971
Earlier work this paper cites.
H. Bock, “Recent advances in parameter identification techniques for ODE,” in Numerical Treatment of Inverse Problems in Differential and Integral Equations , P. Deuflhard and E. Hairer, Eds. Boston: Birkhäuser, 1983, pp. 95–121
1983
Earlier work this paper cites.
R. Fletcher, Practical Methods of Optimization , 2nd ed. Chichester: Wiley, 1987
1987
Earlier work this paper cites.
C. Watkins, “Learning from delayed rewards,” Ph.D. dissertation, King’s College, Cambridge, 1989
1989
Earlier work this paper cites.
F. Y. Wang and I. T. Cameron, “Control studies on a model evaporation process — constrained state driving with conventional and higher relative degree systems,” Journal of Process Control , vol. 4, pp. 59–75, 1994
1994
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Introduction to Reinforcement Learning , 1st ed. Cambridge, MA, USA: MIT Press, 1998
1998
Earlier work this paper cites.
I. Kolmanovsky and E. Gilbert, “Theory and computation of disturbance invariant sets for discrete-time linear systems,” Math. Probl. Eng. , vol. 4, no. 4, pp. 317–367, 1998
1998
Earlier work this paper cites.
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” in Proceedings of the 12th International Conference on Neural Information Processing Systems , ser. NIPS’99. Cambridge, MA, USA: MIT Press, 1999, pp. 1057–1063
1999
Earlier work this paper cites.
P. Scokaert and J. Rawlings, “Feasibility Issues in Linear Model Predictive Control,” AIChE Journal , vol. 45, no. 8, pp. 1649–1659, 1999
1999
Earlier work this paper cites.
L. Chisci, J. Rossiter, and G. Zappa, “Systems with persistent disturbances: predictive control with restricted constraints,” Automatica , vol. 37, pp. 1019–1028, 2001
2001
Earlier work this paper cites.
C. Büskens and H. Maurer, Online Optimization of Large Scale Systems . Berlin, Heidelberg: Springer Berlin Heidelberg, 2001, ch. Sensitivity Analysis and Real-Time Optimization of Parametric Nonlinear Programming Problems, pp. 3–16
2001
Earlier work this paper cites.
J. Nocedal and S. Wright, Numerical Optimization , 2nd ed., ser. Springer Series in Operations Research and Financial Engineering. Springer, 2006
2006
Earlier work this paper cites.
C. Sonntag, O. Stursberg, and S. Engell, “Dynamic Optimization of an Industrial Evaporator using Graph Search with Embedded Nonlinear Programming,” in Proc. 2nd IFAC Conf. on Analysis and Design of Hybrid Systems (ADHS) , 2006, pp. 211–216
2006
Earlier work this paper cites.
P. Abbeel, A. Coates, M. Quigley, and A. Y. Ng, “An application of reinforcement learning to aerobatic helicopter flight,” in In Advances in Neural Information Processing Systems 19 . MIT Press, 2007, p. 2007
2007
Earlier work this paper cites.
F. L. Lewis and D. Vrabie, “Reinforcement learning and adaptive dynamic programming for feedback control,” IEEE Circuits and Systems Magazine , vol. 9, no. 3, pp. 32–50, 2009
2009
Earlier work this paper cites.
M. Diehl, R. Amrit, and J. Rawlings, “A Lyapunov Function for Economic Optimizing Model Predictive Control,” IEEE Trans. of Automatic Control , vol. 56, no. 3, pp. 703–707, March 2011
2011
Earlier work this paper cites.
R. Amrit, J. Rawlings, and D. Angeli, “Economic optimization using model predictive control with a terminal cost,” Annual Reviews in Control , vol. 35, pp. 178–186, 2011
2011
Cited alongside, same era.
R. Amrit, J. B. Rawlings, and L. T. Biegler, “Optimizing process economics online using model predictive control,” Computers & Chemical Engineering , vol. 58, pp. 334 – 343, 2013
2011
Cited alongside, same era.
S. Wang, W. Chaovalitwongse, and R. Babuska, “Machine learning algorithms in bipedal robot control,” Trans. Sys. Man Cyber Part C , vol. 42, no. 5, pp. 728–743, Sep. 2012
2012
Cited alongside, same era.
F. L. Lewis, D. Vrabie, and K. G. Vamvoudakis, “Reinforcement learning and feedback control: Using natural decision methods to design optimal adaptive controllers,” IEEE Control Systems , vol. 32, no. 6, pp. 76–105, 2012
2012
Cited alongside, same era.
M. Zanon, S. Gros, and M. Diehl, “A Tracking MPC Formulation that is Locally Equivalent to Economic MPC,” Journal of Process Control , 2016
2016
Later among the works it cites.
F. Berkenkamp, M. Turchetta, A. Schoellig, and A. Krause, “Safe Model-based Reinforcement Learning with Stability Guarantees,” in Advances in Neural Information Processing Systems 30 , I. Guyon, U. V. Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, Eds. Curran Associates, Inc., 2017, pp. 908–918
2017
Later among the works it cites.
2018
Later among the works it cites.
T. Pham, G. De Magistris, and R. Tachibana, “Optlayer - practical constrained optimization for deep reinforcement learning in the real world,” in 2018 IEEE International Conference on Robotics and Automation (ICRA) , May 2018, pp. 6236–6243
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. B. Rawlings and D. Q. Mayne, Model Predictive Control: Theory and Design . Nob Hill, 2012, ch. Postface to model predictive control: theory and design
2012
Cited alongside, same era.
J. F. J. Garcia, “A comprehensive survey on safe reinforcement learning,” Journal of Machine Learning Research , vol. 16, pp. 1437–1480, 2013
2013
Cited alongside, same era.
A. Aswani, H. Gonzalez, S. S. Sastry, and C. Tomlin, “Provably safe and robust learning-based model predictive control,” Automatica , vol. 49, no. 5, pp. 1216 – 1226, 2013
2013
Cited alongside, same era.
L. Grüne, “Economic receding horizon control without terminal constraints,” Automatica , vol. 49, pp. 725–734, 2013
2013
Cited alongside, same era.
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller, “Deterministic policy gradient algorithms,” in Proceedings of the 31st International Conference on Machine Learning , ser. ICML’14, 2014, pp. I–387–I–395
2014
Cited alongside, same era.
D. Q. Mayne, “Model predictive control: Recent developments and future promise,” Automatica , vol. 50, no. 12, pp. 2967 – 2986, 2014
2014
Cited alongside, same era.
M. Zanon, S. Gros, and M. Diehl, “Indefinite Linear MPC and Approximated Economic MPC for Nonlinear Systems,” Journal of Process Control , vol. 24, pp. 1273–1281, 2014
2014
Cited alongside, same era.
D. Mayne, “Robust and stochastic mpc: Are we going in the right direction?” IFAC-PapersOnLine , vol. 48, no. 23, pp. 1 – 8, 2015, 5th IFAC Conference on Nonlinear Model Predictive Control NMPC 2015
2015
Cited alongside, same era.
2018
Later among the works it cites.
S. Gros and M. Zanon, “Data-Driven Economic NMPC using Reinforcement Learning,” IEEE Transactions on Automatic Control , 2018, (in press)
2018
Later among the works it cites.
T. Koller, F. Berkenkamp, M. Turchetta, and A. Krause, “Learning-based Model Predictive Control for Safe Exploration and Reinforcement Learning,” 2018, published on Arxiv
2018
Later among the works it cites.
R. Murray and M. Palladino, “A model for system uncertainty in reinforcement learning,” Systems & Control Letters , vol. 122, pp. 24 – 31, 2018
2018
Later among the works it cites.
B. Amos, I. D. J. Rodriguez, J. Sacks, B. Boots, and J. Z. Kolter, “Differentiable mpc for end-to-end planning and control,” in Proceedings of the 32Nd International Conference on Neural Information Processing Systems , ser. NIPS’18. USA: Curran Associates Inc., 2018, pp. 8299–8310
2018
Later among the works it cites.
M. Zanon and T. Faulwasser, “Economic MPC without terminal constraints: Gradient-correcting end penalties enforce asymptotic stability,” Journal of Process Control , vol. 63, pp. 1 – 14, 2018
2018
Later among the works it cites.
T. Faulwasser and M. Zanon, “Asymptotic Stability of Economic NMPC: The Importance of Adjoints,” in Proceedings of the IFAC Nonlinear Model Predictive Control Conference , 2018
2018
Later among the works it cites.
L. Grüne and R. Guglielmi, “Turnpike properties and strict dissipativity for discrete time linear quadratic optimal control problems,” SIAM Journal on Control and Optimization , vol. 56, no. 2, pp. 1282–1302, 2018
2018
Later among the works it cites.
M. Zanon, S. Gros, and A. Bemporad, “Practical Reinforcement Learning of Stabilizing Economic MPC,” in Proceedings of the European Control Conference , 2019, (accepted)
2019
Closest in time.
S. Dean, S. Tu, N. Matni, and B. Recht, “Safely Learning to Control the Constrained Linear Quadratic Regulator,” in 2019 American Control Conference (ACC) , July 2019, pp. 5582–5588
2019
Closest in time.
S. Gros and M. Zanon, “Safe Reinforcement Learning Based on Robust MPC and Policy Gradient Methods,” IEEE Transactions on Automatic Control (submitted) , 2019
2019
Closest in time.
S. Gros, M. Zanon, and A. Bemporad, “Safe Reinforcement Learning via Projection on a Safe Set: How to Achieve Optimality?” in 21st IFAC World Congress , 2020, (submitted)
2020
Closest in time.