Fetching the paper…
Reading the bibliography…
Safety is a critical concern when deploying reinforcement learning agents for realistic tasks.
E. Altman, “Constrained markov decision processes with total cost criteria: Lagrangian approach and dual linear program,” Mathematical methods of operations research , vol. 48, no. 3, pp. 387–417, 1998
1998
Earlier work this paper cites.
E. Altman, Constrained Markov decision processes . CRC Press, 1999, vol. 7
1999
Earlier work this paper cites.
A. Nilim and L. Ghaoui, “Robust markov decision problems with uncertain transition matrices,” Advances in Neural Information Processing Systems , 2003
2003
Earlier work this paper cites.
A. Wächter and L. T. Biegler, “On the implementation of an interior-point filter line-search algorithm for large-scale nonlinear programming,” Mathematical programming , vol. 106, no. 1, pp. 25–57, 2006
2006
Earlier work this paper cites.
2013
Earlier work this paper cites.
M. P. Deisenroth, D. Fox, and C. E. Rasmussen, “Gaussian processes for data-efficient learning in robotics and control,” IEEE transactions on pattern analysis and machine intelligence , vol. 37, no. 2, pp. 408–423, 2013
2013
Earlier work this paper cites.
J. R. Pati, “Modeling, identification and control of cart-pole system,” Ph.D. dissertation, 2014
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
J. Garcıa and F. Fernández, “A comprehensive survey on safe reinforcement learning,” Journal of Machine Learning Research , vol. 16, no. 1, pp. 1437–1480, 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
A. Tamar, Y. Glassner, and S. Mannor, “Optimizing the cvar via sampling,” in Proceedings of the Twenty-Ninth AAAI Conference on Artificial Intelligence , 2015, pp. 2993–2999
2015
Earlier work this paper cites.
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al. , “Mastering the game of go with deep neural networks and tree search,” nature , vol. 529, no. 7587, pp. 484–489, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
C. J. Ostafew, A. P. Schoellig, and T. D. Barfoot, “Robust constrained learning-based nmpc enabling reliable mobile robot path tracking,” The International Journal of Robotics Research , vol. 35, no. 13, pp. 1547–1563, 2016
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel, “Domain randomization for transferring deep neural networks from simulation to the real world,” in 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2017, pp. 23–30
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2019
Later among the works it cites.
L. Hewing, J. Kabzan, and M. N. Zeilinger, “Cautious model predictive control using gaussian process regression,” IEEE Transactions on Control Systems Technology , 2019
2019
Later among the works it cites.
W. C. Cheung, D. Simchi-Levi, and R. Zhu, “Non-stationary reinforcement learning: The blessing of (more) optimism,” Available at SSRN 3397818 , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
K. Chua, R. Calandra, R. McAllister, and S. Levine, “Deep reinforcement learning in a handful of trials using probabilistic dynamics models,” in Advances in Neural Information Processing Systems , 2018, pp. 4754–4765
2018
Cited alongside, same era.
T.-H. Pham, G. De Magistris, and R. Tachibana, “Optlayer-practical constrained optimization for deep reinforcement learning in the real world,” in 2018 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2018, pp. 6236–6243
2018
Cited alongside, same era.
Y. Chow, O. Nachum, E. Duenez-Guzman, and M. Ghavamzadeh, “A lyapunov-based approach to safe reinforcement learning,” in Advances in neural information processing systems , 2018, pp. 8092–8101
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
T. Koller, F. Berkenkamp, M. Turchetta, and A. Krause, “Learning-based model predictive control for safe exploration,” in 2018 IEEE Conference on Decision and Control (CDC) . IEEE, 2018, pp. 6059–6066
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2019
Later among the works it cites.
2019
Later among the works it cites.
2019
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
A. Wachi and Y. Sui, “Safe reinforcement learning in constrained markov decision processes,” in International Conference on Machine Learning . PMLR, 2020, pp. 9797–9806
2020
Later among the works it cites.
W. Ding, B. Chen, M. Xu, and D. Zhao, “Learning to collide: An adaptive safety-critical scenarios generating method,” 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
Z. Erickson, V. Gangaram, A. Kapusta, C. K. Liu, and C. C. Kemp, “Assistive gym: A physics simulation framework for assistive robotics,” IEEE International Conference on Robotics and Automation (ICRA) , 2020
2020
Later among the works it cites.
T.-Y. Yang, J. Rosca, K. Narasimhan, and P. J. Ramadge, “Projection-based constrained policy optimization.” in ICLR , 2020
2020
Later among the works it cites.