Fetching the paper…
Reading the bibliography…
Safe reinforcement learning is extremely challenging--not only must the agent explore an unknown environment, it must do so while ensuring no safety constraint violations.
Constrained Markov decision processes
Eitan Altman · 1999
Earlier work this paper cites.
Dynamic programming and optimal control: Vol. 1
Dimitri P Bertsekas et al · 2000
Earlier work this paper cites.
Applications of Markov decision processes in communication networks
Eitan Altman · 2002
Earlier work this paper cites.
An actor-critic algorithm for constrained Markov decision processes
Vivek S Borkar · 2005
Earlier work this paper cites.
Empirical Bernstein bounds and sample-variance penalization
Andreas Maurer and Massimiliano Pontil · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
An online actor–critic algorithm with function approximation for constrained Markov decision processes
Shalabh Bhatnagar and K Lakshmanan · 2012
Earlier work this paper cites.
Strategic planning under uncertainties via constrained Markov decision processes
Xu Chu Ding, Alessandro Pinto, and Amit Surana · 2013
Earlier work this paper cites.
Online learning in episodic Markovian decision processes by relative entropy policy search
Alexander Zimin and Gergely Neu · 2013
Earlier work this paper cites.
Trading safety versus performance: Rapid deployment of robotic swarms with robust performance constraints
Yin-Lam Chow, Marco Pavone, Brian M Sadler, and Stefano Carpin · 2015
Earlier work this paper cites.
Sample complexity of episodic fixed-horizon reinforcement learning
Christoph Dann and Emma Brunskill · 2015
Earlier work this paper cites.
Introduction to online convex optimization
Elad Hazan · 2016
Earlier work this paper cites.
Constrained Policy Optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Earlier work this paper cites.
Risk-constrained reinforcement learning with percentile risk criteria
Yinlam Chow, Mohammad Ghavamzadeh, Lucas Janson, and Marco Pavone · 2017
Earlier work this paper cites.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Earlier work this paper cites.
Online convex optimization with time-varying constraints
Michael J Neely and Hao Yu · 2017
Earlier work this paper cites.
Online convex optimization with stochastic constraints
Hao Yu, Michael J Neely, and Xiaohan Wei · 2017
Cited alongside, same era.
Throughput optimal decentralized scheduling of multihop networks with end-to-end deadline constraints: Unreliable links
Rahul Singh and PR Kumar · 2018
Cited alongside, same era.
Decentralized control via dynamic stochastic prices: The independent system operator problem
Rahul Singh, PR Kumar, and Le Xie · 2018
Cited alongside, same era.
Reward constrained policy optimization
Chen Tessler, Daniel J Mankowitz, and Shie Mannor · 2018
Cited alongside, same era.
Online convex optimization for cumulative constraints
Jianjun Yuan and Andrew Lamperski · 2018
Cited alongside, same era.
Linear stochastic bandits under safety constraints
Sanae Amani, Mahnoosh Alizadeh, and Christos Thrampoulidis · 2019
Learning in Markov decision processes under constraints
Rahul Singh, Abhishek Gupta, and Ness B Shroff · 2020
Later among the works it cites.
First order constrained optimization in policy space
Yiming Zhang, Quan Vuong, and Keith Ross · 2020
Later among the works it cites.
Constrained upper confidence reinforcement learning
Liyuan Zheng and Lillian Ratliff · 2020
Later among the works it cites.
Provably efficient safe exploration via primal-dual policy optimization
Dongsheng Ding, Xiaohan Wei, Zhuoran Yang, Zhaoran Wang, and Mihailo Jovanovic · 2021
Closest in time.
Learning with safety constraints: Sample complexity of reinforcement learning for constrained MDPs
Aria HasanzadeZonuzy, Archana Bura, Dileep Kalathil, and Srinivas Shakkottai · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Constrained EV charging scheduling based on safe deep reinforcement learning
Hepeng Li, Zhiqiang Wan, and Haibo He · 2019
Cited alongside, same era.
Cautious regret minimization: Online optimization with long-term budget constraints
Nikolaos Liakopoulos, Apostolos Destounis, Georgios Paschos, Thrasyvoulos Spyropoulos, and Panayotis Mertikopoulos · 2019
Cited alongside, same era.
Constrained reinforcement learning has zero duality gap
Santiago Paternain, Luiz Chamon, Miguel Calvo-Fullana, and Alejandro Ribeiro · 2019
Cited alongside, same era.
Safe convex learning under uncertain constraints
Ilnura Usmanova, Andreas Krause, and Maryam Kamgarpour · 2019
Cited alongside, same era.
Projection-based constrained policy optimization
Tsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, and Peter J Ramadge · 2019
Cited alongside, same era.
Natural policy gradient primal-dual method for constrained Markov decision processes
Dongsheng Ding, Kaiqing Zhang, Tamer Basar, and Mihailo Jovanovic · 2020
Cited alongside, same era.
Krishna C Kalagarla, Rahul Jain, and Pierluigi Nuzzo · 2021
Closest in time.
Fast global convergence of policy optimization for constrained mdps
Tao Liu, Ruida Zhou, Dileep Kalathil, PR Kumar, and Chao Tian · 2021
Closest in time.
Learning Policies with Zero or Bounded Constraint Violation for Constrained MDPs
Tao Liu, Ruida Zhou, Dileep Kalathil, PR Kumar, and Chao Tian · 2021
Closest in time.
An efficient pessimistic-optimistic algorithm for stochastic linear bandits with general constraints
Xin Liu, Bin Li, Pengyi Shi, and Lei Ying · 2021
Closest in time.
Stochastic bandits with linear constraints
Aldo Pacchiano, Mohammad Ghavamzadeh, Peter Bartlett, and Heinrich Jiang · 2021
Closest in time.
Alwayssafe: Reinforcement learning without safety constraint violations during training
Thiago D Simão, Nils Jansen, and Matthijs TJ Spaan · 2021
Closest in time.
Doubly pessimistic algorithms for strictly safe off-policy optimization
Sanae Amani and Lin F. Yang · 2022
Closest in time.
Safe online convex optimization with unknown linear safety constraints
Sapana Chaudhary and Dileep Kalathil · 2022
Closest in time.
A provably-efficient model-free algorithm for infinite-horizon average-reward constrained Markov decision processes
Honghao Wei, Xin Liu, and Lei Ying · 2022
Closest in time.
Triple-q: A model-free algorithm for constrained reinforcement learning with sublinear regret and zero constraint violation
Honghao Wei, Xin Liu, and Lei Ying · 2022
Closest in time.
Anchor-changing regularized natural policy gradient for multi-objective reinforcement learning
Ruida Zhou, Tao Liu, Dileep Kalathil, PR Kumar, and Chao Tian · 2022
Closest in time.