Fetching the paper…
Reading the bibliography…
Safe reinforcement learning (RL) trains a policy to maximize the task reward while satisfying safety constraints.
A theorem on contraction mappings
Amram Meir and Emmett Keeler · 1969
Earlier work this paper cites.
Constrained markov decision processes with total cost criteria: Lagrangian approach and dual linear program
Eitan Altman · 1998
Earlier work this paper cites.
Introduction to reinforcement learning
Richard S Sutton, Andrew G Barto, et al · 1998
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham Machandranath Kakade · 2003
Earlier work this paper cites.
An online actor–critic algorithm with function approximation for constrained markov decision processes
Shalabh Bhatnagar and K Lakshmanan · 2012
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Javier Garcıa and Fernando Fernández · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
Safe and efficient off-policy reinforcement learning
Rémi Munos, Tom Stepleton, Anna Harutyunyan, and Marc G Bellemare · 2016
Earlier work this paper cites.
Improved techniques for training gans
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen · 2016
Earlier work this paper cites.
Accessorize to a crime: Real and stealthy attacks on state-of-the-art face recognition
Mahmood Sharif, Sruti Bhagavatula, Lujo Bauer, and Michael K Reiter · 2016
Earlier work this paper cites.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Earlier work this paper cites.
Towards principled methods for training generative adversarial networks
Martin Arjovsky and Léon Bottou · 2017
Earlier work this paper cites.
Safe model-based reinforcement learning with stability guarantees
Felix Berkenkamp, Matteo Turchetta, Angela P Schoellig, and Andreas Krause · 2017
Earlier work this paper cites.
Risk-constrained reinforcement learning with percentile risk criteria
Yinlam Chow, Mohammad Ghavamzadeh, Lucas Janson, and Marco Pavone · 2017
Earlier work this paper cites.
Adversarial attacks on neural network policies
Sandy Huang, Nicolas Papernot, Ian Goodfellow, Yan Duan, and Pieter Abbeel · 2017
Earlier work this paper cites.
Delving into adversarial attacks on deep policies
Jernej Kos and Dawn Song · 2017
Earlier work this paper cites.
Tactics of adversarial attack on deep reinforcement learning agents
Yen-Chen Lin, Zhang-Wei Hong, Yuan-Hong Liao, Meng-Li Shih, Ming-Yu Liu, and Min Sun · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2017
Earlier work this paper cites.
Robust deep reinforcement learning with adversarial attacks
Anay Pattanaik, Zhenyi Tang, Shuijing Liu, Gautham Bommannan, and Girish Chowdhary · 2017
Earlier work this paper cites.
Robust adversarial reinforcement learning
Lerrel Pinto, James Davidson, Rahul Sukthankar, and Abhinav Gupta · 2017
Earlier work this paper cites.
Trial without error: Towards safe reinforcement learning via human intervention
William Saunders, Girish Sastry, Andreas Stuhlmueller, and Owain Evans · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel · 2017
Earlier work this paper cites.
Threat of adversarial attacks on deep learning in computer vision: A survey
Naveed Akhtar and Ajmal Mian · 2018
Cited alongside, same era.
Safe reinforcement learning via shielding
Mohammed Alshiekh, Roderick Bloem, Rüdiger Ehlers, Bettina Könighofer, Scott Niekum, and Ufuk Topcu · 2018
Cited alongside, same era.
Safe exploration in continuous action spaces
Gal Dalal, Krishnamurthy Dvijotham, Matej Vecerik, Todd Hester, Cosmin Paduraru, and Yuval Tassa · 2018
Cited alongside, same era.
On the effectiveness of interval bound propagation for training verifiably robust models
Sven Gowal, Krishnamurthy Dvijotham, Robert Stanforth, Rudy Bunel, Chongli Qin, Jonathan Uesato, Relja Arandjelovic, Timothy Mann, and Pushmeet Kohli · 2018
Cited alongside, same era.
Accelerated primal-dual policy optimization for safe reinforcement learning
Safe learning in robotics: From learning-based control to safe reinforcement learning
Lukas Brunke, Melissa Greeff, Adam W Hall, Zhaocong Yuan, Siqi Zhou, Jacopo Panerati, and Angela P Schoellig · 2021
Later among the works it cites.
Context-aware safe reinforcement learning for non-stationary environments
Baiming Chen, Zuxin Liu, Jiacheng Zhu, Mengdi Xu, Wenhao Ding, Liang Li, and Ding Zhao · 2021
Later among the works it cites.
Maximum entropy rl (provably) solves some robust rl problems
Benjamin Eysenbach and Sergey Levine · 2021
Later among the works it cites.
Multi-agent constrained policy optimisation
Shangding Gu, Jakub Grudzien Kuba, Munning Wen, Ruiqing Chen, Ziyan Wang, Zheng Tian, Jun Wang, Alois Knoll, and Yaodong Yang · 2021
Later among the works it cites.
Learning barrier certificates: Towards safe reinforcement learning with zero training-time violations
Yuping Luo and Tengyu Ma · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Qingkai Liang, Fanyu Que, and Eytan Modiano · 2018
Cited alongside, same era.
Domain randomization for simulation-based policy optimization with transferability assessment
Fabio Muratore, Felix Treede, Michael Gienger, and Jan Peters · 2018
Cited alongside, same era.
Learning abstract options
Matthew Riemer, Miao Liu, and Gerald Tesauro · 2018
Cited alongside, same era.
Reward constrained policy optimization
Chen Tessler, Daniel J Mankowitz, and Shie Mannor · 2018
Cited alongside, same era.
Lyapunov-based safe policy optimization for continuous control
Yinlam Chow, Ofir Nachum, Aleksandra Faust, Edgar Duenez-Guzman, and Mohammad Ghavamzadeh · 2019
Cited alongside, same era.
Constrained reinforcement learning has zero duality gap
Santiago Paternain, Luiz FO Chamon, Miguel Calvo-Fullana, and Alejandro Ribeiro · 2019
Cited alongside, same era.
A taxonomy and survey of attacks against machine learning
Nikolaos Pitropakis, Emmanouil Panaousis, Thanassis Giannetsos, Eleftherios Anastasiadis, and George Loukas · 2019
Cited alongside, same era.
Benchmarking safe exploration in deep reinforcement learning
Alex Ray, Joshua Achiam, and Dario Amodei · 2019
Cited alongside, same era.
Later among the works it cites.
Adversarial machine learning in image classification: A survey toward the defender’s perspective
Gabriel Resende Machado, Eugênio Silva, and Ronaldo Ribeiro Goldschmidt · 2021
Later among the works it cites.
Who is the strongest enemy? towards optimal and efficient evasion attacks in deep rl
Yanchao Sun, Ruijie Zheng, Yongyuan Liang, and Furong Huang · 2021
Later among the works it cites.
Recovery rl: Safe reinforcement learning with learned recovery zones
Brijen Thananjeyan, Ashwin Balakrishna, Suraj Nair, Michael Luo, Krishnan Srinivasan, Minho Hwang, Joseph E Gonzalez, Julian Ibarz, Chelsea Finn, and Ken Goldberg · 2021
Later among the works it cites.
Wcsac: Worst-case soft actor critic for safety-constrained reinforcement learning
Qisong Yang, Thiago D Simão, Simon H Tindemans, and Matthijs TJ Spaan · 2021
Later among the works it cites.
Robust reinforcement learning on state observations with learned optimal adversary
Huan Zhang, Hongge Chen, Duane Boning, and Cho-Jui Hsieh · 2021
Later among the works it cites.
Model-free safe control for zero-violation reinforcement learning
Weiye Zhao, Tairan He, and Changliu Liu · 2021
Later among the works it cites.
Constrained policy optimization via bayesian world models
Yarden As, Ilnura Usmanova, Sebastian Curi, and Andreas Krause · 2022
Closest in time.
Saac: Safe reinforcement learning as an adversarial game of actor-critics
Yannis Flet-Berliac and Debabrota Basu · 2022
Closest in time.
Bullet-safety-gym: Aframework for constrained reinforcement learning
Sven Gronauer · 2022
Closest in time.
A review of safe reinforcement learning: Methods, theory and applications
Shangding Gu, Long Yang, Yali Du, Guang Chen, Florian Walter, Jun Wang, Yaodong Yang, and Alois Knoll · 2022
Closest in time.
Robust reinforcement learning as a stackelberg game via adaptively-regularized adversarial training
Peide Huang, Mengdi Xu, Fei Fang, and Ding Zhao · 2022
Closest in time.
Efficient off-policy safe reinforcement learning using trust region conditional value at risk
Dohyeong Kim and Songhwai Oh · 2022
Closest in time.
Deep reinforcement learning policies learn shared adversarial features across mdps
Ezgi Korkmaz · 2022
Closest in time.
Efficient adversarial training without attacking: Worst-case-aware robust reinforcement learning
Yongyuan Liang, Yanchao Sun, Ruijie Zheng, and Furong Huang · 2022
Closest in time.
Constrained variational policy optimization for safe reinforcement learning
Zuxin Liu, Zhepeng Cen, Vladislav Isenbaev, Wei Liu, Steven Wu, Bo Li, and Ding Zhao · 2022
Closest in time.
Robust reinforcement learning: A review of foundations and recent advances
Janosch Moos, Kay Hansel, Hany Abdulsamad, Svenja Stark, Debora Clever, and Jan Peters · 2022
Closest in time.
Towards safe reinforcement learning with a safety editor policy
Haonan Yu, Wei Xu, and Haichao Zhang · 2022
Closest in time.
Adversarial robust deep reinforcement learning requires redefining robustness
Ezgi Korkmaz · 2023
Closest in time.