Fetching the paper…
Reading the bibliography…
We study the problem of computing an optimal policy of an infinite-horizon discounted constrained Markov decision process (constrained MDP).
Non-cooperative games
John Nash Jr · 1951
Earlier work this paper cites.
Studies in linear and non-linear programming
Kenneth J. Arrow, Leonid Hurwicz, and Hirofumi Uzawa · 1958
Earlier work this paper cites.
On general minimax theorems
Maurice Sion · 1958
Earlier work this paper cites.
Functional approximations and dynamic programming
Richard Bellman and Stuart Dreyfus · 1959
Earlier work this paper cites.
The extragradient method for finding saddle points and other problems
Galina M Korpelevich · 1976
Earlier work this paper cites.
A modification of the Arrow-Hurwicz method for search of saddle points
Leonid Denisovich Popov · 1980
Earlier work this paper cites.
Optimal policies for controlled Markov chains with a constraint
Frederick J Beutler and Keith W Ross · 1985
Earlier work this paper cites.
A convex analytic approach to Markov decision processes
Vivek S Borkar · 1988
Earlier work this paper cites.
Randomized and past-dependent policies for Markov decision processes with multiple constraints
Keith W Ross · 1989
Earlier work this paper cites.
Control of random sequences in problems with constraints
Alexey Piunovskiy · 1994
Earlier work this paper cites.
On linear convergence of iterative methods for the variational inequality problem
Paul Tseng · 1995
Earlier work this paper cites.
Constrained discounted dynamic programming
Eugene A Feinberg and Adam Shwartz · 1996
Earlier work this paper cites.
Primal-dual interior-point methods
Stephen J Wright · 1997
Earlier work this paper cites.
Constrained Markov decision processes
Eitan Altman · 1999
Earlier work this paper cites.
Self learning control of constrained Markov chains – a gradient approach
F Vazquez Abad, Vikram Krishnamurthy, Katerine Martin, and Irina Baltcheva · 2002
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Martin Zinkevich · 2003
Earlier work this paper cites.
An actor-critic algorithm for constrained Markov decision processes
Vivek S Borkar · 2005
Earlier work this paper cites.
Numerical solution of saddle point problems
Michele Benzi, Gene H Golub, and Jörg Liesen · 2005
Earlier work this paper cites.
Lagrange multiplier approach to variational problems and applications
Kazufumi Ito and Karl Kunisch · 2008
Earlier work this paper cites.
Subgradient methods for saddle-point problems
Angelia Nedić and Asuman Ozdaglar · 2009
Earlier work this paper cites.
Optimizing debt collections using constrained reinforcement learning
Naoki Abe, Prem Melville, Cezar Pendus, Chandan K Reddy, David L Jensen, Vince P Thomas, James J Bennett, Gary F Anderson, Brent R Cooley, Melissa Kowalczyk, et al · 2010
Earlier work this paper cites.
On the acceleration of augmented Lagrangian method for linearly constrained optimization
Bingsheng He and Xiaoming Yuan · 2010
Earlier work this paper cites.
An online actor–critic algorithm with function approximation for constrained Markov decision processes
Shalabh Bhatnagar and K Lakshmanan · 2012
Earlier work this paper cites.
Simon Lacoste-Julien, Mark Schmidt, and Francis Bach · 2012
Earlier work this paper cites.
Online learning with predictable sequences
Alexander Rakhlin and Karthik Sridharan · 2013
Earlier work this paper cites.
Reduction of constrained maxima to saddle-point problems
Kenneth J Arrow and Leonid Hurwicz · 2013
Earlier work this paper cites.
Constrained optimization and Lagrange multiplier methods
Dimitri P Bertsekas · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Javier Garcıa and Fernando Fernández · 2015
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Sébastien Bubeck · 2015
Earlier work this paper cites.
Projected reflected gradient methods for monotone variational inequalities
Yu Malitsky · 2015
Earlier work this paper cites.
A universal primal-dual convex optimization framework
Alp Yurtsever, Quoc Tran Dinh, and Volkan Cevher · 2015
Earlier work this paper cites.
On non-ergodic convergence rate of Douglas–Rachford alternating direction method of multipliers
Bingsheng He and Xiaoming Yuan · 2015
Earlier work this paper cites.
Continuity of optimal solution functions and their conditions on objective functions
Yasushi Terazono and Ayumu Matani · 2015
Earlier work this paper cites.
Fast projection onto the simplex and the l 1 l_{1} ball
Laurent Condat · 2016
Earlier work this paper cites.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Earlier work this paper cites.
Risk-constrained reinforcement learning with percentile risk criteria
Yinlam Chow, Mohammad Ghavamzadeh, Lucas Janson, and Marco Pavone · 2017
Earlier work this paper cites.
Non-ergodic alternating proximal augmented Lagrangian algorithms with optimal rates
Quoc Tran Dinh · 2018
Earlier work this paper cites.
Augmented Lagrangian-based decomposition methods with non-ergodic optimal rates
Quoc Tran-Dinh and Yuzixuan Zhu · 2018
Earlier work this paper cites.
Reward constrained policy optimization
Chen Tessler, Daniel J Mankowitz, and Shie Mannor · 2019
Earlier work this paper cites.
On the convergence of single-call stochastic extra-gradient methods
Yu-Guan Hsieh, Franck Iutzeler, Jérôme Malick, and Panayotis Mertikopoulos · 2019
Earlier work this paper cites.
Sample-optimal parametric Q-learning using linearly additive features
Lin Yang and Mengdi Wang · 2019
Earlier work this paper cites.
Last-iterate convergence: Zero-sum games and constrained min-max optimization
C Daskalakis and Ioannis Panageas · 2019
Earlier work this paper cites.
Convergent policy optimization for safe reinforcement learning
Ming Yu, Zhuoran Yang, Mladen Kolar, and Zhaoran Wang · 2019
Earlier work this paper cites.
Constrained reinforcement learning has zero duality gap
Santiago Paternain, Luiz Chamon, Miguel Calvo-Fullana, and Alejandro Ribeiro · 2019
Cited alongside, same era.
On the nonergodic convergence rate of an inexact augmented Lagrangian framework for composite convex programming
Ya-Feng Liu, Xin Liu, and Shiqian Ma · 2019
Cited alongside, same era.
Two-player games for efficient non-convex constrained optimization
Andrew Cotter, Heinrich Jiang, and Karthik Sridharan · 2019
Cited alongside, same era.
Natural policy gradient primal-dual method for constrained Markov decision processes
Dongsheng Ding, Kaiqing Zhang, Tamer Başar, and Mihailo R Jovanović · 2020
Cited alongside, same era.
Responsive safety in reinforcement learning by PID Lagrangian methods
Adam Stooke, Joshua Achiam, and Pieter Abbeel · 2020
Cited alongside, same era.
Linear last-iterate convergence in constrained saddle-point optimization
A review of safe reinforcement learning: Methods, theory and applications
Shangding Gu, Long Yang, Yali Du, Guang Chen, Florian Walter, Jun Wang, Yaodong Yang, and Alois Knoll · 2022
Later among the works it cites.
Towards painless policy optimization for constrained MDPs
Arushi Jain, Sharan Vaswani, Reza Babanezhad, Csaba Szepesvari, and Doina Precup · 2022
Later among the works it cites.
Finite-time complexity of online primal-dual natural actor-critic algorithm for constrained Markov decision processes
Sihan Zeng, Thinh T Doan, and Justin Romberg · 2022
Later among the works it cites.
A dual approach to constrained Markov decision processes with entropy regularization
Donghao Ying, Yuhao Ding, and Javad Lavaei · 2022
Later among the works it cites.
Constrained reinforcement learning via dissipative saddle flow dynamics
Tianqi Zheng, Pengcheng You, and Enrique Mallada · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen-Yu Wei, Chung-Wei Lee, Mengxiao Zhang, and Haipeng Luo · 2020
Cited alongside, same era.
Efficiently solving MDPs with stochastic mirror descent
Yujia Jin and Aaron Sidford · 2020
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Cited alongside, same era.
Neural policy gradient methods: Global optimality and rates of convergence
Lingxiao Wang, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2020
Cited alongside, same era.
Projection-based constrained policy optimization
Tsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, and Peter J Ramadge · 2020
Cited alongside, same era.
First order constrained optimization in policy space
Yiming Zhang, Quan Vuong, and Keith Ross · 2020
Cited alongside, same era.
IPO: Interior-point policy optimization under constraints
Yongshuai Liu, Jiaxin Ding, and Xin Liu · 2020
Cited alongside, same era.
Later among the works it cites.
Yang Cai, Argyris Oikonomou, and Weiqiang Zheng · 2022
Later among the works it cites.
Near-optimal sample complexity bounds for constrained MDPs
Sharan Vaswani, Lin Yang, and Csaba Szepesvári · 2022
Later among the works it cites.
Achieving zero constraint violation for constrained reinforcement learning via primal-dual approach
Qinbo Bai, Amrit Singh Bedi, Mridul Agarwal, Alec Koppel, and Vaneet Aggarwal · 2022
Later among the works it cites.
Dongsheng Ding, Kaiqing Zhang, Jiali Duan, Tamer Başar, and Mihailo R Jovanović · 2022
Later among the works it cites.
Fast global convergence of natural policy gradient methods with entropy regularization
Shicong Cen, Chen Cheng, Yuxin Chen, Yuting Wei, and Yuejie Chi · 2022
Later among the works it cites.
Policy gradient primal-dual mirror descent for constrained MDPs with large state spaces
Dongsheng Ding and Mihailo R Jovanović · 2022
Later among the works it cites.
Triple-Q: A model-free algorithm for constrained reinforcement learning with sublinear regret and zero constraint violation
Honghao Wei, Xin Liu, and Lei Ying · 2022
Later among the works it cites.
Linear convergence of natural policy gradient methods with log-linear policies
Rui Yuan, Simon Shaolei Du, Robert M Gower, Alessandro Lazaric, and Lin Xiao · 2022
Later among the works it cites.
On the convergence rates of policy gradient methods
Lin Xiao · 2022
Later among the works it cites.
Convergence and optimality of policy gradient primal-dual method for constrained Markov decision processes
Dongsheng Ding, Kaiqing Zhang, Tamer Başar, and Mihailo R Jovanović · 2022
Later among the works it cites.
Anchor-changing regularized natural policy gradient for multi-objective reinforcement learning
Ruida Zhou, Tao Liu, Dileep Kalathil, Panganamala Kumar, and Chao Tian · 2022
Later among the works it cites.
CUP: A conservative update policy algorithm for safe reinforcement learning, 2022
Long Yang, Yu Zhang, Jiaming Ji, Juntao Dai, and Weidong Zhang · 2022
Later among the works it cites.
A simple reward-free approach to constrained reinforcement learning
Sobhan Miryoosefi and Chi Jin · 2022
Later among the works it cites.
A provably-efficient model-free algorithm for infinite-horizon average-reward constrained Markov decision processes
Honghao Wei, Xin Liu, and Lei Ying · 2022
Later among the works it cites.
Learning infinite-horizon average-reward Markov decision process with constraints
Liyu Chen, Rahul Jain, and Haipeng Luo · 2022
Later among the works it cites.
DOPE: Doubly optimistic and pessimistic exploration for safe reinforcement learning
Archana Bura, Aria HasanzadeZonuzy, Dileep Kalathil, Srinivas Shakkottai, and Jean-Francois Chamberland · 2022
Later among the works it cites.
Provably efficient model-free constrained RL with linear function approximation
Arnob Ghosh, Xingyu Zhou, and Ness Shroff · 2022
Later among the works it cites.
Learning in constrained Markov decision processes
Rahul Singh, Abhishek Gupta, and Ness B Shroff · 2022
Later among the works it cites.
A policy gradient approach for finite horizon constrained Markov decision processes
Soumyajit Guin and Shalabh Bhatnagar · 2022
Later among the works it cites.
Yang Cai, Argyris Oikonomou, and Weiqiang Zheng · 2022
Later among the works it cites.
Faster Lagrangian-based methods in convex optimization
Shoham Sabach and Marc Teboulle · 2022
Later among the works it cites.
Offline reinforcement learning with realizability and single-policy concentrability
Wenhao Zhan, Baihe Huang, Audrey Huang, Nan Jiang, and Jason Lee · 2022
Later among the works it cites.
Safe-state enhancement method for autonomous driving via direct hierarchical reinforcement learning
Ziqing Gu, Lingping Gao, Haitong Ma, Shengbo Eben Li, Sifa Zheng, Wei Jing, and Junbo Chen · 2023
Closest in time.
State augmented constrained reinforcement learning: Overcoming the limitations of learning with rewards
Miguel Calvo-Fullana, Santiago Paternain, Luiz FO Chamon, and Alejandro Ribeiro · 2023
Closest in time.
Algorithm for constrained Markov decision process with linear convergence
Egor Gladin, Maksim Lavrik-Karmazin, Karina Zainullina, Varvara Rudenko, Alexander Gasnikov, and Martin Takac · 2023
Closest in time.
Ted Moskovitz, Brendan O’Donoghue, Vivek Veeriah, Sebastian Flennerhag, Satinder Singh, and Tom Zahavy · 2023
Closest in time.
Can we find Nash equilibria at a linear rate in Markov games?
Zhuoqing Song, Jason D Lee, and Zhuoran Yang · 2023
Closest in time.
Constrained MDPs and the reward hypothesis
Csaba Szepesvári · 2023
Closest in time.
Achieving zero constraint violation for constrained reinforcement learning via conservative natural policy gradient primal-dual algorithm
Qinbo Bai, Amrit Singh Bedi, and Vaneet Aggarwal · 2023
Closest in time.
Provably efficient model-free algorithms for non-stationary CMDPs
Honghao Wei, Arnob Ghosh, Ness Shroff, Lei Ying, and Xingyu Zhou · 2023
Closest in time.
A best-of-both-worlds algorithm for constrained MDPs with long-term constraints
Jacopo Germano, Francesco Emanuele Stradi, Gianmarco Genalti, Matteo Castiglioni, Alberto Marchesi, and Nicola Gatti · 2023
Closest in time.
Safe posterior sampling for constrained MDPs with bounded constraint violation
Krishna C Kalagarla, Rahul Jain, and Pierluigi Nuzzo · 2023
Closest in time.
Policy mirror descent for regularized reinforcement learning: A generalized framework with linear convergence
Wenhao Zhan, Shicong Cen, Baihe Huang, Yuxin Chen, Jason D Lee, and Yuejie Chi · 2023
Closest in time.
Policy mirror descent for reinforcement learning: Linear convergence, new sampling complexity, and generalized problem classes
Guanghui Lan · 2023
Closest in time.
Extragradient-type methods with O ( 1 / k ) (1/k) -convergence rates for co-hypomonotone inclusions
Quoc Tran-Dinh · 2023
Closest in time.
Tao Zhang, Yong Xia, and Shiru Li · 2023
Closest in time.
Achieving zero constraint violation for concave utility constrained reinforcement learning via primal-dual approach
Qinbo Bai, Amrit Singh Bedi, Mridul Agarwal, Alec Koppel, and Vaneet Aggarwal · 2023
Closest in time.
Policy-based primal-dual methods for convex constrained Markov decision processes
Donghao Ying, Mengzi Amy Guo, Yuhao Ding, Javad Lavaei, and Zuo-Jun Shen · 2023
Closest in time.
Introduction to online optimization/learning: Lecture 1
Haipeng Luo · 2023
Closest in time.