Fetching the paper…
Reading the bibliography…
For all its successes, Reinforcement Learning (RL) still struggles to deliver formal guarantees on the closed-loop behavior of the learned policy.
Safe Reinforcement Learning Using Robust MPC
Zanon, M. and Gros (2019) · 1906
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R.S., McAllester, D., Singh, S., and Mansour, Y. (1999) · 1999
Earlier work this paper cites.
Dynamic Programming and Optimal Control , volume 2
Bertsekas, D. (2007) · 2007
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
J. Garcia, J.F. (2013) · 2013
Cited alongside, same era.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M. (2014) · 2014
Cited alongside, same era.
Reinforcement learning: An introduction. Second Edition
Sutton, R.S. and Barto, A.G. (2018) · 2018
Cited alongside, same era.
Probabilistic model predictive safety certification for learning-based control
Wabersich, K., Hewing, L., Carron, A., and Zeilinger, M. (2019) · 2019
Later among the works it cites.
Safe Reinforcement Learning Based on Robust MPC and Policy Gradient Methods
Gros, S. and Zanon, M. (2020) · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…