Fetching the paper…
Reading the bibliography…
In this paper, we investigate a novel safe reinforcement learning problem with step-wise violation constraints.
Probability inequalities for sums of bounded random variables
Wassily Hoeffding · 1994
Earlier work this paper cites.
Constrained reinforcement learning from intrinsic and extrinsic rewards
Eiji Uchibe and Kenji Doya · 2007
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Peter Auer, Thomas Jaksch, and Ronald Ortner · 2008
Earlier work this paper cites.
On lower bounds for regret in reinforcement learning
Ian Osband and Benjamin Van Roy · 2016
Earlier work this paper cites.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Earlier work this paper cites.
Safe model-based reinforcement learning with stability guarantees
Felix Berkenkamp, Matteo Turchetta, Angela Schoellig, and Andreas Krause · 2017
Earlier work this paper cites.
Safe reinforcement learning via shielding
Mohammed Alshiekh, Roderick Bloem, Rüdiger Ehlers, Bettina Könighofer, Scott Niekum, and Ufuk Topcu · 2018
Earlier work this paper cites.
A lyapunov-based approach to safe reinforcement learning
Yinlam Chow, Ofir Nachum, Edgar Duenez-Guzman, and Mohammad Ghavamzadeh · 2018
Earlier work this paper cites.
Safe exploration in continuous action spaces
Gal Dalal, Krishnamurthy Dvijotham, Matej Vecerik, Todd Hester, Cosmin Paduraru, and Yuval Tassa · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Earlier work this paper cites.
Reward constrained policy optimization
Chen Tessler, Daniel J Mankowitz, and Shie Mannor · 2018
Earlier work this paper cites.
Openspiel: A framework for reinforcement learning in games
Marc Lanctot, Edward Lockhart, Jean-Baptiste Lespiau, Vinicius Zambaldi, Satyaki Upadhyay, Julien Pérolat, Sriram Srinivasan, Finbarr Timbers, Karl Tuyls, Shayegan Omidshafiei, et al · 2019
Earlier work this paper cites.
Batch policy learning under constraints
Hoang Le, Cameron Voloshin, and Yisong Yue · 2019
Earlier work this paper cites.
Non-asymptotic gap-dependent regret bounds for tabular mdps
Max Simchowitz and Kevin G Jamieson · 2019
Earlier work this paper cites.
Convergent policy optimization for safe reinforcement learning
Ming Yu, Zhuoran Yang, Mladen Kolar, and Zhaoran Wang · 2019
Cited alongside, same era.
Natural policy gradient primal-dual method for constrained markov decision processes
Dongsheng Ding, Kaiqing Zhang, Tamer Basar, and Mihailo Jovanovic · 2020
Cited alongside, same era.
Exploration-exploitation in constrained mdps
Yonathan Efroni, Shie Mannor, and Matteo Pirotta · 2020
Cited alongside, same era.
Reward-free exploration for reinforcement learning
Chi Jin, Akshay Krishnamurthy, Max Simchowitz, and Tiancheng Yu · 2020
Cited alongside, same era.
Ipo: Interior-point policy optimization under constraints
Yongshuai Liu, Jiaxin Ding, and Xin Liu · 2020
Cited alongside, same era.
Upper confidence primal-dual reinforcement learning for cmdp with adversarial loss
A sample-efficient algorithm for episodic finite-horizon mdp with constraints
Krishna C Kalagarla, Rahul Jain, and Pierluigi Nuzzo · 2021
Later among the works it cites.
Adaptive reward-free exploration
Emilie Kaufmann, Pierre Ménard, Omar Darwiche Domingues, Anders Jonsson, Edouard Leurent, and Michal Valko · 2021
Later among the works it cites.
Fast active learning for pure exploration in reinforcement learning
Pierre Ménard, Omar Darwiche Domingues, Anders Jonsson, Emilie Kaufmann, Edouard Leurent, and Michal Valko · 2021
Later among the works it cites.
Alwayssafe: Reinforcement learning without safety constraint violations during training
Thiago D Simão, Nils Jansen, and Matthijs TJ Spaan · 2021
Later among the works it cites.
Safe reinforcement learning by imagining the near future
Garrett Thomas, Yuping Luo, and Tengyu Ma · 2021
Later among the works it cites.
Provably efficient algorithms for multi-objective competitive rl
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shuang Qiu, Xiaohan Wei, Zhuoran Yang, Jieping Ye, and Zhaoran Wang · 2020
Cited alongside, same era.
Learning in markov decision processes under constraints
Rahul Singh, Abhishek Gupta, and Ness B Shroff · 2020
Cited alongside, same era.
Responsive safety in reinforcement learning by pid lagrangian methods
Adam Stooke, Joshua Achiam, and Pieter Abbeel · 2020
Cited alongside, same era.
Safe reinforcement learning via curriculum induction
Matteo Turchetta, Andrey Kolobov, Shital Shah, Andreas Krause, and Alekh Agarwal · 2020
Cited alongside, same era.
Safe reinforcement learning in constrained markov decision processes
Akifumi Wachi and Yanan Sui · 2020
Cited alongside, same era.
Sim-to-real transfer in deep reinforcement learning for robotics: a survey
Wenshuai Zhao, Jorge Peña Queralta, and Tomi Westerlund · 2020
Cited alongside, same era.
Safe reinforcement learning with linear function approximation
Sanae Amani, Christos Thrampoulidis, and Lin Yang · 2021
Cited alongside, same era.
Tiancheng Yu, Yi Tian, Jingzhao Zhang, and Suvrit Sra · 2021
Later among the works it cites.
Reinforcement learning based recommender systems: A survey
M Mehdi Afsar, Trafford Crump, and Behrouz Far · 2022
Later among the works it cites.
Safe exploration incurs nearly no additional sample complexity for reward-free rl
Ruiquan Huang, Jing Yang, and Yingbin Liang · 2022
Later among the works it cites.
A simple reward-free approach to constrained reinforcement learning
Sobhan Miryoosefi and Chi Jin · 2022
Later among the works it cites.
Enhancing safe exploration using safety state augmentation
Aivar Sootla, Alexander Cowen-Rivers, Jun Wang, and Haitham Bou Ammar · 2022
Later among the works it cites.
Yixuan Wang, Simon Sinong Zhan, Ruochen Jiao, Zhilu Wang, Wanxin Jin, Zhuoran Yang, Zhaoran Wang, Chao Huang, and Qi Zhu · 2022
Later among the works it cites.
Triple-q: A model-free algorithm for constrained reinforcement learning with sublinear regret and zero constraint violation
Honghao Wei, Xin Liu, and Lei Ying · 2022
Later among the works it cites.
Reachability constrained reinforcement learning
Dongjie Yu, Haitong Ma, Shengbo Li, and Jianyu Chen · 2022
Later among the works it cites.
A near-optimal algorithm for safe reinforcement learning under instantaneous hard constraints
Ming Shi, Yingbin Liang, and Ness Shroff · 2023
Closest in time.