Fetching the paper…
Reading the bibliography…
We study the sequential decision making problem of maximizing the expected total reward while satisfying a constraint on the expected total utility.
On the global optimum convergence of momentum-based policy gradient
Yuhao Ding, Junzi Zhang, and Javad Lavaei · 1934
Earlier work this paper cites.
Studies in Linear and Non-linear Programming
Kenneth J Arrow · 1958
Earlier work this paper cites.
Linear and Nonlinear Programming , volume 2
David G Luenberger and Yinyu Ye · 1984
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E Schapire · 1997
Earlier work this paper cites.
Optimal Control of Random Sequences in Problems with Constraints
A.B. Piunovskiy · 1997
Earlier work this paper cites.
Constrained Markov Decision Processes , volume 7
Eitan Altman · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Self learning control of constrained Markov chains – a gradient approach
Felisa Vázquez Abad, Vikram Krishnamurthy, Katerine Martin, and Irina Baltcheva · 2002
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham M Kakade and John Langford · 2002
Earlier work this paper cites.
Policy gradient stochastic approximation algorithms for adaptive control of constrained time varying Markov decision processes
Felisa Vázquez Abad and Vikram Krishnamurthy · 2003
Earlier work this paper cites.
An actor-critic algorithm for constrained Markov decision processes
Vivek S Borkar · 2005
Earlier work this paper cites.
Prediction, Learning, and Games
Nicolo Cesa-Bianchi and Gábor Lugosi · 2006
Earlier work this paper cites.
Constrained reinforcement learning from intrinsic and extrinsic rewards
Eiji Uchibe and Kenji Doya · 2007
Earlier work this paper cites.
Nonlinear Programming: Second Edition
Dimitri P Bertsekas · 2008
Earlier work this paper cites.
Online Markov decision processes
Eyal Even-Dar, Sham M Kakade, and Yishay Mansour · 2009
Earlier work this paper cites.
Optimizing debt collections using constrained reinforcement learning
Naoki Abe, Prem Melville, Cezar Pendus, Chandan K Reddy, David L Jensen, Vince P Thomas, James J Bennett, Gary F Anderson, Brent R Cooley, Melissa Kowalczyk, et al · 2010
Earlier work this paper cites.
An online actor–critic algorithm with function approximation for constrained Markov decision processes
Shalabh Bhatnagar and K Lakshmanan · 2012
Earlier work this paper cites.
Simon Lacoste-Julien, Mark Schmidt, and Francis Bach · 2012
Earlier work this paper cites.
Trading regret for efficiency: online convex optimization with long term constraints
Mehrdad Mahdavi, Rong Jin, and Tianbao Yang · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Non-strongly-convex smooth stochastic approximation with convergence rate O ( 1 / n ) {O}(1/n)
Francis Bach and Eric Moulines · 2013
Earlier work this paper cites.
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
Ohad Shamir and Tong Zhang · 2013
Earlier work this paper cites.
Constrained Optimization and Lagrange Multiplier Methods
Dimitri P Bertsekas · 2014
Earlier work this paper cites.
Understanding Machine Learning: From Theory to Algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
Chance-constrained dynamic programming with application to risk-aware robotic space exploration
Masahiro Ono, Marco Pavone, Yoshiaki Kuwata, and J Balaram · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Continuity of optimal solution functions and their conditions on objective functions
Yasushi Terazono and Ayumu Matani · 2015
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Earlier work this paper cites.
First-order Methods in Optimization , volume 25
Amir Beck · 2017
Earlier work this paper cites.
Risk-constrained reinforcement learning with percentile risk criteria
Yinlam Chow, Mohammad Ghavamzadeh, Lucas Janson, and Marco Pavone · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Online convex optimization with stochastic constraints
Hao Yu, Michael Neely, and Xiaohan Wei · 2017
Cited alongside, same era.
A Lyapunov-based approach to safe reinforcement learning
Yinlam Chow, Ofir Nachum, Edgar Duenez-Guzman, and Mohammad Ghavamzadeh · 2018
Cited alongside, same era.
Global convergence of policy gradient methods for the linear quadratic regulator
Maryam Fazel, Rong Ge, Sham M Kakade, and Mehran Mesbahi · 2018
Cited alongside, same era.
Safe reinforcement learning with linear function approximation
Sanae Amani, Christos Thrampoulidis, and Lin F Yang · 2021
Later among the works it cites.
On the linear convergence of policy gradient methods for finite MDPs
Jalaj Bhandari and Daniel Russo · 2021
Later among the works it cites.
Provably efficient safe exploration via primal-dual policy optimization
Dongsheng Ding, Xiaohan Wei, Zhuoran Yang, Zhaoran Wang, and Mihailo Jovanović · 2021
Later among the works it cites.
On the linear convergence of random search for discrete-time LQR
Hesameddin Mohammadi, Mahdi Soltanolkotabi, and Mihailo R Jovanović · 2021
Later among the works it cites.
CRPO: A new approach for safe reinforcement learning with convergence guarantee
Tengyu Xu, Yingbin Liang, and Guanghui Lan · 2021
Later among the works it cites.
Provably efficient algorithms for multi-objective competitive RL
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jaime F Fisac, Anayo K Akametalu, Melanie N Zeilinger, Shahab Kaynama, Jeremy Gillula, and Claire J Tomlin · 2018
Cited alongside, same era.
Accelerated primal-dual policy optimization for safe reinforcement learning
Qingkai Liang, Fanyu Que, and Eytan Modiano · 2018
Cited alongside, same era.
Stagewise safe Bayesian optimization with Gaussian processes
Yanan Sui, Joel Burdick, Yisong Yue, et al · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Online convex optimization for cumulative constraints
Jianjun Yuan and Andrew Lamperski · 2018
Cited alongside, same era.
Neural trust region/proximal policy optimization attains globally optimal policy
Boyi Liu, Qi Cai, Zhuoran Yang, and Zhaoran Wang · 2019
Cited alongside, same era.
Global exponential convergence of gradient methods over the nonconvex landscape of the linear quadratic regulator
Hesameddin Mohammadi, Armin Zare, Mahdi Soltanolkotabi, and Mihailo R Jovanović · 2019
Cited alongside, same era.
Tiancheng Yu, Yi Tian, Jingzhao Zhang, and Suvrit Sra · 2021
Later among the works it cites.
Achieving zero constraint violation for constrained reinforcement learning via primal-dual approach
Qinbo Bai, Amrit Singh Bedi, Mridul Agarwal, Alec Koppel, and Vaneet Aggarwal · 2022
Closest in time.
Fast global convergence of natural policy gradient methods with entropy regularization
Shicong Cen, Chen Cheng, Yuxin Chen, Yuting Wei, and Yuejie Chi · 2022
Closest in time.
Policy gradient primal-dual mirror descent for constrained MDPs with large state spaces
Dongsheng Ding and Mihailo R Jovanović · 2022
Closest in time.
Convergence and optimality of policy gradient primal-dual method for constrained markov decision processes
Dongsheng Ding, Kaiqing Zhang, Tamer Başar, and Mihailo R Jovanović · 2022
Closest in time.
Towards painless policy optimization for constrained MDPs
Arushi Jain, Sharan Vaswani, Reza Babanezhad, Csaba Szepesvari, and Doina Precup · 2022
Closest in time.
On linear and super-linear convergence of natural policy gradient algorithm
Sajad Khodadadian, Prakirt Raj Jhunjhunwala, Sushil Mahavir Varma, and Siva Theja Maguluri · 2022
Closest in time.
A simple reward-free approach to constrained reinforcement learning
Sobhan Miryoosefi and Chi Jin · 2022
Closest in time.
Convergence and sample complexity of gradient methods for the model-free linear-quadratic regulator problem
Hesameddin Mohammadi, Armin Zare, Mahdi Soltanolkotabi, and Mihailo R Jovanović · 2022
Closest in time.
Safe policies for reinforcement learning via primal-dual methods
Santiago Paternain, Miguel Calvo-Fullana, Luiz FO Chamon, and Alejandro Ribeiro · 2022
Closest in time.
Learning in constrained Markov decision processes
Rahul Singh, Abhishek Gupta, and Ness B Shroff · 2022
Closest in time.
A provably-efficient model-free algorithm for infinite-horizon average-reward constrained Markov decision processes
Honghao Wei, Xin Liu, and Lei Ying · 2022
Closest in time.
Constrained update projection approach to safe policy optimization
Long Yang, Jiaming Ji, Juntao Dai, Linrui Zhang, Binbin Zhou, Pengfei Li, Yaodong Yang, and Gang Pan · 2022
Closest in time.
A dual approach to constrained markov decision processes with entropy regularization
Donghao Ying, Yuhao Ding, and Javad Lavaei · 2022
Closest in time.
Last-iterate convergent policy gradient primal-dual methods for constrained MDPs
Dongsheng Ding, Chen-Yu Wei, Kaiqing Zhang, and Alejandro Ribeiro · 2023
Closest in time.
Stochastic policy gradient methods: Improved sample complexity for fisher-non-degenerate policies
Ilyas Fatkhullin, Anas Barakat, Anastasia Kireeva, and Niao He · 2023
Closest in time.
Reload: Reinforcement learning with optimistic ascent-descent for last-iterate convergence in constrained MDPs
Ted Moskovitz, Brendan O’Donoghue, Vivek Veeriah, Sebastian Flennerhag, Satinder Singh, and Tom Zahavy · 2023
Closest in time.
Global optimality guarantees for policy gradient methods
Jalaj Bhandari and Daniel Russo · 2024
Closest in time.
Resilient constrained reinforcement learning
Dongsheng Ding, Zhengyan Huan, and Alejandro Ribeiro · 2024
Closest in time.
A primal-dual-critic algorithm for offline constrained reinforcement learning
Kihyuk Hong, Yuhang Li, and Ambuj Tewari · 2024
Closest in time.
Omnisafe: An infrastructure for accelerating safe reinforcement learning research
Jiaming Ji, Jiayi Zhou, Borong Zhang, Juntao Dai, Xuehai Pan, Ruiyang Sun, Weidong Huang, Yiran Geng, Mickel Liu, and Yaodong Yang · 2024
Closest in time.
Faster algorithm and sharper analysis for constrained Markov decision process
Tianjiao Li, Ziwei Guan, Shaofeng Zou, Tengyu Xu, Yingbin Liang, and Guanghui Lan · 2024
Closest in time.
Sample-efficient constrained reinforcement learning with general parameterization
Washim U Mondal and Vaneet Aggarwal · 2024
Closest in time.
Last-iterate global convergence of policy gradients for constrained reinforcement learning
Alessandro Montenegro, Marco Mussi, Matteo Papini, and Alberto Maria Metelli · 2024
Closest in time.
Truly no-regret learning in constrained MDPs
Adrian Müller, Pragnya Alatur, Volkan Cevher, Giorgia Ramponi, and Niao He · 2024
Closest in time.
Adversarially trained weighted actor-critic for safe offline reinforcement learning
Honghao Wei, Xiyue Peng, Arnob Ghosh, and Xin Liu · 2024
Closest in time.
Distributionally robust constrained reinforcement learning under strong duality
Zhengfei Zhang, Kishan Panaganti, Laixi Shi, Yanan Sui, Adam Wierman, and Yisong Yue · 2024
Closest in time.
Deterministic policy gradient primal-dual methods for continuous-space constrained MDPs
Sergio Rozada, Dongsheng Ding, Antonio G Marques, and Alejandro Ribeiro · 2025
Closest in time.