Fetching the paper…
Reading the bibliography…
In safe Reinforcement Learning (RL), safety cost is typically defined as a function dependent on the immediate state and actions.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E., and Williams, R. J · 1986
Earlier work this paper cites.
Constrained markov decision processes with total cost criteria: Lagrangian approach and dual linear program
Altman, E · 1998
Earlier work this paper cites.
Introduction to reinforcement learning. vol. 135, 1998
Sutton, R. S., Barto, A. G., et al · 1998
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y., Russell, S., et al · 2000
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
On the properties of neural machine translation: Encoder–decoder approaches
Cho, K., van Merriënboer, B., Bahdanau, D., and Bengio, Y · 2014
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Garcıa, J. and Fernández, F · 2015
Earlier work this paper cites.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P · 2017
Earlier work this paper cites.
Safe reinforcement learning via shielding
Alshiekh, M., Bloem, R., Ehlers, R., Könighofer, B., Niekum, S., and Topcu, U · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
Penalizing side effects using stepwise relative reachability
Krakovna, V., Orseau, L., Kumar, R., Martic, M., and Legg, S · 2018
Earlier work this paper cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Levine, S · 2018
Earlier work this paper cites.
Reward constrained policy optimization
Tessler, C., Mankowitz, D. J., and Mannor, S · 2018
Cited alongside, same era.
Minimax-regret querying on side effects for safe optimality in factored markov decision processes
Zhang, S., Durfee, E. H., and Singh, S · 2018
Cited alongside, same era.
Benchmarking safe exploration in deep reinforcement learning
Ray, A., Achiam, J., and Amodei, D · 2019
Cited alongside, same era.
Preferences implicit in the state of the world
Shah, R. and Krasheninnikov, D · 2019
Cited alongside, same era.
Deep reinforcement learning for power system applications: An overview
Zhang, Z., Zhang, D., and Qiu, R. C · 2019
Cited alongside, same era.
Safe reinforcement learning using probabilistic shields
Jansen, N., Könighofer, B., Junges, S., Serban, A., and Bloem, R · 2020
Escaping from zero gradient: Revisiting action-constrained reinforcement learning via frank-wolfe policy optimization
Lin, J.-L., Hung, W., Yang, S.-H., Hsieh, P.-C., and Liu, X · 2021
Later among the works it cites.
Inverse constrained reinforcement learning
Malik, S., Anwar, U., Aghasi, A., and Ahmed, A · 2021
Later among the works it cites.
Avoiding negative side effects due to incomplete knowledge of AI systems
Saisubramanian, S., Zilberstein, S., and Kamar, E · 2021
Later among the works it cites.
Recovery rl: Safe reinforcement learning with learned recovery zones
Thananjeyan, B., Balakrishna, A., Nair, S., Luo, M., Srinivasan, K., Hwang, M., Gonzalez, J. E., Ibarz, J., Finn, C., and Goldberg, K · 2021
Later among the works it cites.
Safe reinforcement learning by imagining the near future
Thomas, G., Luo, Y., and Ma, T · 2021
Later among the works it cites.
Learning soft constraints from constrained expert demonstrations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Responsive safety in reinforcement learning by pid lagrangian methods
Stooke, A., Achiam, J., and Abbeel, P · 2020
Cited alongside, same era.
Energy-efficient heating control for smart buildings with deep reinforcement learning
Gupta, A., Badr, Y., Negahban, A., and Qiu, R. G · 2021
Cited alongside, same era.
Learning to walk in the real world with minimal human effort
Ha, S., Xu, P., Tan, Z., Levine, S., and Tan, J · 2021
Cited alongside, same era.
Compositional reinforcement learning from logical specifications
Jothimurugan, K., Bansal, S., Bastani, O., and Alur, R · 2021
Cited alongside, same era.
Deep reinforcement learning for autonomous driving: A survey
Kiran, B. R., Sobh, I., Talpaert, V., Mannion, P., Al Sallab, A. A., Yogamani, S., and Pérez, P · 2021
Cited alongside, same era.
Avoiding side effects in complex environments
Turner, A., Ratzlaff, N., and Tadepalli, P
Cited in the paper.
Gaurav, A., Rezaee, K., Liu, G., and Poupart, P · 2022
Later among the works it cites.
Bullet-safety-gym: A framework for constrained reinforcement learning
Gronauer, S · 2022
Later among the works it cites.
Benchmarking constraint inference in inverse reinforcement learning
Liu, G., Luo, Y., Gaurav, A., Rezaee, K., and Poupart, P · 2022
Later among the works it cites.
Avoiding negative side effects of autonomous systems in the open world
Saisubramanian, S., Kamar, E., and Zilberstein, S · 2022
Later among the works it cites.
pyRDDLGym: From RDDL to Gym Environments
Taitler, A., Gimelfarb, M., Gopalakrishnan, S., Mladenov, M., Liu, X., and Sanner, S · 2022
Later among the works it cites.
Benchmarking actor-critic deep reinforcement learning algorithms for robotics control with action constraints
Kasaura, K., Miura, S., Kozuno, T., Yonetani, R., Hoshino, K., and Hosoe, Y · 2023
Later among the works it cites.