Fetching the paper…
Reading the bibliography…
We address the problem of finding the optimal policy of a constrained Markov decision process (CMDP) using a gradient descent-based algorithm.
Constrained Markov decision processes
E. Altman · 1999
Earlier work this paper cites.
A natural policy gradient
S. M. Kakade · 2001
Earlier work this paper cites.
Concentration inequalities: A nonasymptotic theory of independence
S. Boucheron, G. Lugosi, and P. Massart · 2013
Earlier work this paper cites.
Constrained optimization and Lagrange multiplier methods
D. P. Bertsekas · 2014
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
S. Bubeck · 2015
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
J. Garcıa and F. Fernández · 2015
Earlier work this paper cites.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz · 2015
Earlier work this paper cites.
Concrete problems in AI safety
D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané · 2016
Earlier work this paper cites.
Openai gym, 2016
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
A cmdp-based approach for energy efficient power allocation in massive mimo systems
P. Li, Y. Jiang, W. Li, F. Zheng, and X. You · 2016
Earlier work this paper cites.
Constrained policy optimization
J. Achiam, D. Held, A. Tamar, and P. Abbeel · 2017
Earlier work this paper cites.
A unified view of entropy-regularized markov decision processes
G. Neu, A. Jonsson, and V. Gómez · 2017
Earlier work this paper cites.
H. Yu and M. J. Neely · 2017
Cited alongside, same era.
A simple parallel algorithm with an O(1/t) convergence rate for general convex programs
H. Yu and M. J. Neely · 2017
Cited alongside, same era.
Accelerated primal-dual policy optimization for safe reinforcement learning
Q. Liang, F. Que, and E. Modiano · 2018
Cited alongside, same era.
Reward constrained policy optimization
C. Tessler, D. J. Mankowitz, and S. Mannor · 2018
Cited alongside, same era.
Safe policies for reinforcement learning via primal-dual methods
S. Paternain, M. Calvo-Fullana, L. F. Chamon, and A. Ribeiro · 2019
Cited alongside, same era.
Y. Xu · 2020
Later among the works it cites.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
A. Agarwal, S. M. Kakade, J. D. Lee, and G. Mahajan · 2021
Closest in time.
Fast global convergence of natural policy gradient methods with entropy regularization
S. Cen, C. Cheng, Y. Chen, Y. Wei, and Y. Chi · 2021
Closest in time.
Provably efficient safe exploration via primal-dual policy optimization
D. Ding, X. Wei, Z. Yang, Z. Wang, and M. Jovanovic · 2021
Closest in time.
Challenges of real-world reinforcement learning: definitions, benchmarks and analysis
G. Dulac-Arnold, N. Levine, D. J. Mankowitz, J. Li, C. Paduraru, S. Gowal, and T. Hester · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Projection-based constrained policy optimization
T.-Y. Yang, J. Rosca, K. Narasimhan, and P. J. Ramadge · 2019
Cited alongside, same era.
Optimality and approximation with policy gradient methods in markov decision processes
A. Agarwal, S. M. Kakade, J. D. Lee, and G. Mahajan · 2020
Cited alongside, same era.
Natural policy gradient primal-dual method for constrained markov decision processes
D. Ding, K. Zhang, T. Basar, and M. R. Jovanovic · 2020
Cited alongside, same era.
Exploration-exploitation in constrained MDPs
Y. Efroni, S. Mannor, and M. Pirotta · 2020
Cited alongside, same era.
On the global convergence rates of softmax policy gradient methods
J. Mei, C. Xiao, C. Szepesvari, and D. Schuurmans · 2020
Cited alongside, same era.
Responsive safety in reinforcement learning by pid lagrangian methods
A. Stooke, J. Achiam, and P. Abbeel · 2020
Cited alongside, same era.
Online primal-dual mirror descent under stochastic constraints
X. Wei, H. Yu, and M. J. Neely · 2020
Cited alongside, same era.
G. Lan · 2021
Closest in time.
Faster algorithm and sharper analysis for constrained markov decision process
T. Li, Z. Guan, S. Zou, T. Xu, Y. Liang, and G. Lan · 2021
Closest in time.
Learning policies with zero or bounded constraint violation for constrained mdps
T. Liu, R. Zhou, D. Kalathil, P. R. Kumar, and C. Tian · 2021
Closest in time.
CRPO: A new approach for safe reinforcement learning with convergence guarantee
T. Xu, Y. Liang, and G. Lan · 2021
Closest in time.
A dual approach to constrained markov decision processes with entropy regularization
D. Ying, Y. Ding, and J. Lavaei · 2021
Closest in time.
W. Zhan, S. Cen, B. Huang, Y. Chen, J. D. Lee, and Y. Chi · 2021
Closest in time.