Fetching the paper…
Reading the bibliography…
Safety comes first in many real-world applications involving autonomous agents.
Value constrained model-free continuous control
Bohez, S.; Abdolmaleki, A.; Neunert, M.; Buchli, J.; Heess, N.; and Hadsell, R. 2019 · 1902
Earlier work this paper cites.
Benchmarking safe exploration in deep reinforcement learning
Ray, A.; Achiam, J.; and Amodei, D. 2019 · 1910
Earlier work this paper cites.
Linear and nonlinear programming , volume 2
Luenberger, D. G.; Ye, Y.; et al. 1984 · 1984
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S.; and Barto, A. G. 1998 · 1998
Earlier work this paper cites.
Constrained Markov decision processes , volume 7
Altman, E. 1999 · 1999
Earlier work this paper cites.
Learning to walk in the real world with minimal human effort
Ha, S.; Xu, P.; Tan, Z.; Levine, S.; and Tan, J. 2020 · 2002
Earlier work this paper cites.
First order constrained optimization in policy space
Zhang, Y.; Vuong, Q.; and Ross, K. W. 2020 · 2002
Earlier work this paper cites.
Learning to be safe: Deep rl with a safety critic
Srinivasan, K.; Eysenbach, B.; Ha, S.; Tan, J.; and Finn, C. 2020 · 2010
Earlier work this paper cites.
Projection-based constrained policy optimization
Yang, T.-Y.; Rosca, J.; Narasimhan, K.; and Ramadge, P. J. 2020 · 2010
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A. A.; Veness, J.; Bellemare, M. G.; Graves, A.; Riedmiller, M.; Fidjeland, A. K.; Ostrovski, G.; et al. 2015 · 2015
Earlier work this paper cites.
Constrained policy optimization
Achiam, J.; Held, D.; Tamar, A.; and Abbeel, P. 2017 · 2017
Earlier work this paper cites.
Safe exploration in continuous action spaces
Dalal, G.; Dvijotham, K.; Vecerik, M.; Hester, T.; Paduraru, C.; and Tassa, Y. 2018 · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Fujimoto, S.; Hoof, H.; and Meger, D. 2018 · 2018
Cited alongside, same era.
Survey on human–robot collaboration in industrial settings: Safety, intuitive interfaces and applications
Villani, V.; Pini, F.; Leali, F.; and Secchi, C. 2018 · 2018
Cited alongside, same era.
Safe exploration and optimization of constrained mdps using gaussian processes
Wachi, A.; Sui, Y.; Yue, Y.; and Ono, M. 2018 · 2018
Cited alongside, same era.
End-to-end safe reinforcement learning through barrier functions for safety-critical continuous control tasks
Cheng, R.; Orosz, G.; Murray, R. M.; and Burdick, J. W. 2019 · 2019
Cited alongside, same era.
Dc3: A learning method for optimization with hard constraints
Donti, P. L.; Rolnick, D.; and Kolter, J. Z. 2021 · 2021
Later among the works it cites.
Learn-to-race: A multimodal control environment for autonomous racing
Herman, J.; Francis, J.; Ganju, S.; Chen, B.; Koul, A.; Gupta, A.; Skabelkin, A.; Zhukov, I.; Kumskoy, M.; and Nyberg, E. 2021 · 2021
Later among the works it cites.
Deep reinforcement learning for autonomous driving: A survey
Kiran, B. R.; Sobh, I.; Talpaert, V.; Mannion, P.; Al Sallab, A. A.; Yogamani, S.; and Pérez, P. 2021 · 2021
Later among the works it cites.
Metadrive: Composing diverse driving scenarios for generalizable reinforcement learning
Li, Q.; Peng, Z.; Xue, Z.; Zhang, Q.; and Zhou, B. 2021 · 2021
Later among the works it cites.
Feasible actor-critic: Constrained reinforcement learning for ensuring statewise safety
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
First-order methods almost always avoid saddle points: The case of vanishing step-sizes
Panageas, I.; Piliouras, G.; and Wang, X. 2019 · 2019
Cited alongside, same era.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Vinyals, O.; Babuschkin, I.; Czarnecki, W. M.; Mathieu, M.; Dudzik, A.; Chung, J.; Choi, D. H.; Powell, R.; Ewalds, T.; Georgiev, P.; et al. 2019 · 2019
Cited alongside, same era.
Ipo: Interior-point policy optimization under constraints
Liu, Y.; Ding, J.; and Liu, X. 2020 · 2020
Cited alongside, same era.
Safe reinforcement learning in constrained Markov decision processes
Wachi, A.; and Sui, Y. 2020 · 2020
Cited alongside, same era.
Safe Autonomous Racing via Approximate Reachability on Ego-vision
Chen, B.; Francis, J.; Nyberg, J. O. E.; and Herbert, S. L. 2021 · 2021
Cited alongside, same era.
Constrained Update Projection Approach to Safe Policy Optimization
Yang, L.; Ji, J.; Dai, J.; Zhang, L.; Zhou, B.; Li, P.; Yang, Y.; and Pan, G. 2022a
Cited in the paper.
Safe Reinforcement Learning for Legged Locomotion
Yang, T.-Y.; Zhang, T.; Luu, L.; Ha, S.; Tan, J.; and Yu, W. 2022b
Cited in the paper.
Ma, H.; Guan, Y.; Li, S. E.; Zhang, X.; Zheng, S.; and Chen, J. 2021 · 2021
Later among the works it cites.
Recovery rl: Safe reinforcement learning with learned recovery zones
Thananjeyan, B.; Balakrishna, A.; Nair, S.; Luo, M.; Srinivasan, K.; Hwang, M.; Gonzalez, J. E.; Ibarz, J.; Finn, C.; and Goldberg, K. 2021 · 2021
Later among the works it cites.
Yuan, Z.; Hall, A. W.; Zhou, S.; Brunke, L.; Greeff, M.; Panerati, J.; and Schoellig, A. P. 2021 · 2021
Later among the works it cites.
Model-free safe control for zero-violation reinforcement learning
Zhao, W.; He, T.; and Liu, C. 2021 · 2021
Later among the works it cites.
Reachability Constrained Reinforcement Learning
Yu, D.; Ma, H.; Li, S.; and Chen, J. 2022 · 2022
Closest in time.
Penalized Proximal Policy Optimization for Safe Reinforcement Learning
Zhang, L.; Shen, L.; Yang, L.; Chen, S.; Wang, X.; Yuan, B.; and Tao, D. 2022 · 2022
Closest in time.