Fetching the paper…
Reading the bibliography…
We consider the safe reinforcement learning (RL) problem of maximizing utility with extremely low constraint violation rates.
Nonlinear Programming
Bertsekas, D. (1999) · 1999
Earlier work this paper cites.
Autonomous helicopter flight via reinforcement learning
Kim, H., Jordan, M., Sastry, S., and Ng, A. (2004) · 2004
Earlier work this paper cites.
An online actor–critic algorithm with function approximation for constrained markov decision processes
Bhatnagar, S. and Lakshmanan, K. (2012) · 2012
Earlier work this paper cites.
A survey of multi-objective sequential decision-making
Roijers, D. M., Vamplew, P., Whiteson, S., and Dazeley, R. (2013) · 2013
Earlier work this paper cites.
Scalarized multi-objective reinforcement learning: Novel design techniques
Van Moffaert, K., Drugan, M. M., and Nowé, A. (2013) · 2013
Earlier work this paper cites.
Multi-objective reinforcement learning using sets of pareto dominating policies
Moffaert, K. V. and Nowé, A. (2014) · 2014
Earlier work this paper cites.
Openai gym
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Earlier work this paper cites.
Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection
Levine, S., Pastor, P., Krizhevsky, A., and Quillen, D. (2016) · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D. (2016) · 2016
Earlier work this paper cites.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P. (2017) · 2017
Earlier work this paper cites.
Optnet: Differentiable optimization as a layer in neural networks
Amos, B. and Kolter, J. Z. (2017) · 2017
Earlier work this paper cites.
Safe model-based reinforcement learning with stability guarantees
Berkenkamp, F., Turchetta, M., Schoellig, A. P., and Krause, A. (2017) · 2017
Earlier work this paper cites.
Improving stochastic policy gradients in continuous control with deep reinforcement learning using the beta distribution
Chou, P.-W., Maturana, D., and Scherer, S. (2017) · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017) · 2017
Earlier work this paper cites.
A lyapunov-based approach to safe reinforcement learning
Chow, Y., Nachum, O., Duenez-Guzman, E., and Ghavamzadeh, M. (2018) · 2018
Earlier work this paper cites.
Safe exploration in continuous action spaces
Dalal, G., Dvijotham, K., Vecerík, M., Hester, T., Paduraru, C., and Tassa, Y. (2018) · 2018
Cited alongside, same era.
Soft actor-critic algorithms and applications
Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., and Levine, S. (2018) · 2018
Cited alongside, same era.
Optlayer - practical constrained optimization for deep reinforcement learning in the real world
Pham, T.-H., Magistris, G. D., and Tachibana, R. (2018) · 2018
Cited alongside, same era.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., Józefowicz, R., Gray, S., Olsson, C., Pachocki, J., Petrov, M., de Oliveira Pinto, H. P., Raiman, J., Salimans, T., Schlatter, J., Schneider, J., Sidor, S., Sutskever, I., Tang, J., Wolski, F., and Zhang, S. (2019) · 2019
Cited alongside, same era.
Value constrained model-free continuous control
Bohez, S., Abdolmaleki, A., Neunert, M., Buchli, J., Heess, N., and Hadsell, R. (2019) · 2019
Projection-based constrained policy optimization
Yang, T.-Y., Rosca, J., Narasimhan, K., and Ramadge, P. J. (2020) · 2020
Later among the works it cites.
First order constrained optimization in policy space
Zhang, Y., Vuong, Q., and Ross, K. W. (2020) · 2020
Later among the works it cites.
Sim-to-real transfer in deep reinforcement learning for robotics: a survey
Zhao, W., Queralta, J. P., and Westerlund, T. (2020) · 2020
Later among the works it cites.
How rl agents behave when their actions are modified
Langlois, E. D. and Everitt, T. (2021) · 2021
Later among the works it cites.
Safe reinforcement learning using robust action governor
Li, Y., Li, N., Tseng, H. E., Girard, A. R., Filev, D., and Kolmanovsky, I. V. (2021) · 2021
Later among the works it cites.
Debiasing a first-order heuristic for approximate bi-level optimization
Likhosherstov, V., Song, X., Choromanski, K., Davis, J., and Weller, A. (2021) · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
End-to-end safe reinforcement learning through barrier functions for safety-critical continuous control tasks
Cheng, R., Orosz, G., Murray, R. M., and Burdick, J. W. (2019) · 2019
Cited alongside, same era.
Lyapunov-based safe policy optimization for continuous control
Chow, Y., Nachum, O., Faust, A., Duenez-Guzman, E., and Ghavamzadeh, M. (2019) · 2019
Cited alongside, same era.
Solving rubik’s cube with a robot hand
OpenAI, Akkaya, I., Andrychowicz, M., Chociej, M., Litwin, M., McGrew, B., Petron, A., Paino, A., Plappert, M., Powell, G., Ribas, R., Schneider, J., Tezak, N., Tworek, J., Welinder, P., Weng, L., Yuan, Q., Zaremba, W., and Zhang, L. (2019) · 2019
Cited alongside, same era.
Benchmarking Safe Exploration in Deep Reinforcement Learning
Ray, A., Achiam, J., and Amodei, D. (2019) · 2019
Cited alongside, same era.
Reward constrained policy optimization
Tessler, C., Mankowitz, D. J., and Mannor, S. (2019) · 2019
Cited alongside, same era.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., Oh, J., Horgan, D., Kroiss, M., Danihelka, I., Huang, A., Sifre, L., Cai, T., Agapiou, J. P., Jaderberg, M., Vezhnevets, A. S., Leblond, R., Pohlen, T., Dalibard, V., Budden, D., Sulsky, Y., Molloy, J., Paine, T. L., Gulcehre, C., Wang, Z., Pfaff, T., Wu, Y., Ring, R., Yogatama, D., Wünsch, D., McKinney, K., Smith, O., Schaul, T., Lillicrap, T. P., Kavukcuoglu, K., Hassabis, D., Apps, C., and Silver, D. (2019) · 2019
Cited alongside, same era.
Safety aware reinforcement learning (sarl)
Miret, S., Majumdar, S., and Wainwright, C. (2020) · 2020
Cited alongside, same era.
Later among the works it cites.
Learning barrier certificates: Towards safe reinforcement learning with zero training-time violations
Luo, Y. and Ma, T. (2021) · 2021
Later among the works it cites.
DESTA: A framework for safe reinforcement learning with markov games of intervention
Mguni, D., Jennings, J., Jafferjee, T., Sootla, A., Yang, Y., Yu, C., Islam, U., Wang, Z., and Wang, J. (2021) · 2021
Later among the works it cites.
Density constrained reinforcement learning
Qin, Z., Chen, Y., and Fan, C. (2021) · 2021
Later among the works it cites.
Recovery RL: safe reinforcement learning with learned recovery zones
Thananjeyan, B., Balakrishna, A., Nair, S., Luo, M., Srinivasan, K., Hwang, M., Gonzalez, J. E., Ibarz, J., Finn, C., and Goldberg, K. (2021) · 2021
Later among the works it cites.
Safe reinforcement learning by imagining the near future
Thomas, G., Luo, Y., and Ma, T. (2021) · 2021
Later among the works it cites.
Constrained policy optimization via bayesian world models
As, Y., Usmanova, I., Curi, S., and Krause, A. (2022) · 2022
Closest in time.
SAAC: safe reinforcement learning as an adversarial game of actor-critics
Flet-Berliac, Y. and Basu, D. (2022) · 2022
Closest in time.
Constrained variational policy optimization for safe reinforcement learning
Liu, Z., Cen, Z., Isenbaev, V., Liu, W., Wu, Z., Li, B., and Zhao, D. (2022) · 2022
Closest in time.