E. Altman, Constrained Markov decision processes . CRC Press, 1999
1999
Earlier work this paper cites.
N. Kohl and P. Stone, “Policy gradient reinforcement learning for fast quadrupedal locomotion,” in IEEE International Conference on Robotics and Automation, 2004. Proceedings , 2004
2004
Earlier work this paper cites.
R. Tedrake, T. W. Zhang, and H. S. Seung, “Stochastic policy gradient reinforcement learning on a simple 3d biped,” in IEEE/RSJ International Conference on Intelligent Robots and Systems , 2004
2004
Earlier work this paper cites.
G. Endo, J. Morimoto, T. Matsubara, J. Nakanishi, and G. Cheng, “Learning cpg sensory feedback with policy gradient for biped locomotion for a full-body humanoid.” 2005
2005
Earlier work this paper cites.
J. Garcia and F. Fernandez, “A comprehensive survey on safe reinforcement learning,” Journal of Machine Learning Research , 2015
2015
Earlier work this paper cites.
J. Schulman, S. Levine, P. Abbeel, M. I. Jordan, and P. Moritz, “Trust region policy optimization,” in Proceedings of the 32nd International Conference on Machine Learning, ICML 2015 . JMLR.org, 2015
2015
Earlier work this paper cites.
S. Junges, N. Jansen, C. Dehnert, U. Topcu, and J.-P. Katoen, “Safety-constrained reinforcement learning for mdps,” in International Conference on Tools and Algorithms for the Construction and Analysis of Systems . Springer, 2016
2016
Earlier work this paper cites.
J. Achiam, D. Held, A. Tamar, and P. Abbeel, “Constrained policy optimization,” in Proceedings of the 34th International Conference on Machine Learning, ICML , 2017
2017
Earlier work this paper cites.
F. Berkenkamp, M. Turchetta, A. P. Schoellig, and A. Krause, “Safe model-based reinforcement learning with stability guarantees,” in Advances in Neural Information Processing Systems 30 , 2017
2017
Earlier work this paper cites.
E. Coumans and Y. Bai, “Pybullet, a python module for physics simulation in robotics, games and machine learning,” 2017
2017
Earlier work this paper cites.
T. Haarnoja, S. Ha, A. Zhou, J. Tan, G. Tucker, and S. Levine, “Learning to walk via deep reinforcement learning,” arXiv preprint arXiv:1812.11103 , 2018
Original
2018
Earlier work this paper cites.
Y. Chow, O. Nachum, E. A. Duéñez-Guzmán, and M. Ghavamzadeh, “A lyapunov-based approach to safe reinforcement learning,” in Advances in Neural Information Processing Systems 31, NeurIPS , 2018
2018
Earlier work this paper cites.
G. Dalal, K. Dvijotham, M. Vecerik, T. Hester, C. Paduraru, and Y. Tassa, “Safe exploration in continuous action spaces,” arXiv preprint arXiv:1801.08757 , 2018
Original
2018
Earlier work this paper cites.
M. Fu and A. Prashanth L, “Risk-sensitive reinforcement learning: A constrained optimization viewpoint,” arXiv preprint arXiv:1810.09126 , 2018
Original
2018
Earlier work this paper cites.
J. F. Fisac, A. K. Akametalu, M. N. Zeilinger, S. Kaynama, J. Gillula, and C. J. Tomlin, “A general safety framework for learning-based control in uncertain robotic systems,” IEEE Transactions on Automatic Control , no. 7, 2018
2018
Earlier work this paper cites.