Fetching the paper…
Reading the bibliography…
Safe reinforcement learning (RL) aims to learn policies that satisfy certain constraints before deploying them to safety-critical applications.
On the lambertw function
Corless, R. M., Gonnet, G. H., Hare, D. E., Jeffrey, D. J., and Knuth, D. E · 1996
Earlier work this paper cites.
Optimization by vector space methods , pp. 213–216
Luenberger, D. G · 1997
Earlier work this paper cites.
Constrained markov decision processes with total cost criteria: Lagrangian approach and dual linear program
Altman, E · 1998
Earlier work this paper cites.
Tutorial on maximum likelihood estimation
Myung, I. J · 2003
Earlier work this paper cites.
On integral probability metrics, \ \backslash phi-divergences and binary classification
Sriperumbudur, B. K., Fukumizu, K., Gretton, A., Schölkopf, B., and Lanckriet, G. R · 2009
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Earlier work this paper cites.
Safe and efficient off-policy reinforcement learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M. G · 2016
Earlier work this paper cites.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P · 2017
Earlier work this paper cites.
Risk-constrained reinforcement learning with percentile risk criteria
Chow, Y., Ghavamzadeh, M., Janson, L., and Pavone, M · 2017
Earlier work this paper cites.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
Trial without error: Towards safe reinforcement learning via human intervention
Saunders, W., Sastry, G., Stuhlmueller, A., and Evans, O · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Cited alongside, same era.
Safe reinforcement learning via shielding
Alshiekh, M., Bloem, R., Ehlers, R., Könighofer, B., Niekum, S., and Topcu, U · 2018
Cited alongside, same era.
A lyapunov-based approach to safe reinforcement learning
Chow, Y., Nachum, O., Duenez-Guzman, E., and Ghavamzadeh, M · 2018
Cited alongside, same era.
Safe exploration in continuous action spaces
Dalal, G., Dvijotham, K., Vecerik, M., Hester, T., Paduraru, C., and Tassa, Y · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Hoof, H., and Meger, D · 2018
Cited alongside, same era.
V-mpo: On-policy maximum a posteriori policy optimization for discrete and continuous control
Song, H. F., Abdolmaleki, A., Springenberg, J. T., Clark, A., Soyer, H., Rae, J. W., Noury, S., Ahuja, A., Liu, S., Tirumala, D., et al · 2019
Later among the works it cites.
Natural policy gradient primal-dual method for constrained markov decision processes
Ding, D., Zhang, K., Basar, T., and Jovanovic, M · 2020
Later among the works it cites.
Constrained model-based reinforcement learning with robust cross-entropy method
Liu, Z., Zhou, H., Chen, B., Zhong, S., Hebert, M., and Zhao, D · 2020
Later among the works it cites.
Responsive safety in reinforcement learning by pid lagrangian methods
Stooke, A., Achiam, J., and Abbeel, P · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Levine, S · 2018
Cited alongside, same era.
Accelerated primal-dual policy optimization for safe reinforcement learning
Liang, Q., Que, F., and Modiano, E · 2018
Cited alongside, same era.
Reward constrained policy optimization
Tessler, C., Mankowitz, D. J., and Mannor, S · 2018
Cited alongside, same era.
Value constrained model-free continuous control
Bohez, S., Abdolmaleki, A., Neunert, M., Buchli, J., Heess, N., and Hadsell, R · 2019
Cited alongside, same era.
End-to-end safe reinforcement learning through barrier functions for safety-critical continuous control tasks
Cheng, R., Orosz, G., Murray, R. M., and Burdick, J. W · 2019
Cited alongside, same era.
Virel: A variational inference framework for reinforcement learning
Fellows, M., Mahajan, A., Rudner, T. G., and Whiteson, S · 2019
Cited alongside, same era.
Yang, T.-Y., Rosca, J., Narasimhan, K., and Ramadge, P. J · 2020
Later among the works it cites.
First order constrained optimization in policy space
Zhang, Y., Vuong, Q., and Ross, K · 2020
Later among the works it cites.
Context-aware safe reinforcement learning for non-stationary environments
Chen, B., Liu, Z., Zhu, J., Xu, M., Ding, W., and Zhao, D · 2021
Later among the works it cites.
Recovery rl: Safe reinforcement learning with learned recovery zones
Thananjeyan, B., Balakrishna, A., Nair, S., Luo, M., Srinivasan, K., Hwang, M., Gonzalez, J. E., Ibarz, J., Finn, C., and Goldberg, K · 2021
Later among the works it cites.
Safe reinforcement learning using advantage-based intervention
Wagener, N., Boots, B., and Cheng, C.-A · 2021
Later among the works it cites.
Crpo: A new approach for safe reinforcement learning with convergence guarantee
Xu, T., Liang, Y., and Lan, G · 2021
Later among the works it cites.
On the properties of kullback-leibler divergence between gaussians
Zhang, Y., Liu, W., Chen, Z., Li, K., and Wang, J · 2021
Later among the works it cites.
Constrained policy optimization via bayesian world models
As, Y., Usmanova, I., Curi, S., and Krause, A · 2022
Closest in time.
Bullet-safety-gym: Aframework for constrained reinforcement learning
Gronauer, S · 2022
Closest in time.