Fetching the paper…
Reading the bibliography…
We study the constrained reinforcement learning problem, in which an agent aims to maximize the expected cumulative reward subject to a constraint on the expected total value of a utility function.
Hitting-time and occupation-time bounds implied by drift analysis with applications
Bruce Hajek · 1982
Earlier work this paper cites.
Complexity analysis of real-time reinforcement learning
Sven Koenig and Reid G Simmons · 1993
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Roman Vershynin · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Yasin Abbasi-Yadkori, Dávid Pál, and Csaba Szepesvári · 2011
Earlier work this paper cites.
On the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar, Rémi Munos, and Bert Kappen · 2012
Earlier work this paper cites.
What doubling tricks can and can’t do for multi-armed bandits
Lilian Besson and Emilie Kaufmann · 2018
Earlier work this paper cites.
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Earlier work this paper cites.
Sample-optimal parametric q-learning using linearly additive features
Lin Yang and Mengdi Wang · 2019
Earlier work this paper cites.
Safe policies for reinforcement learning via primal-dual methods
Santiago Paternain, Miguel Calvo-Fullana, Luiz FO Chamon, and Alejandro Ribeiro · 2019
Earlier work this paper cites.
Reinforcement learning with dynamic boltzmann softmax updates
Ling Pan, Qingpeng Cai, Qi Meng, Wei Chen, Longbo Huang, and Tie-Yan Liu · 2019
Earlier work this paper cites.
Exploration-exploitation in constrained mdps
Yonathan Efroni, Shie Mannor, and Matteo Pirotta · 2020
Earlier work this paper cites.
Learning in markov decision processes under constraints
Rahul Singh, Abhishek Gupta, and Ness B Shroff · 2020
Earlier work this paper cites.
Constrained episodic reinforcement learning in concave-convex and knapsack settings
Kianté Brantley, Miroslav Dudik, Thodoris Lykouris, Sobhan Miryoosefi, Max Simchowitz, Aleksandrs Slivkins, and Wen Sun · 2020
Cited alongside, same era.
Constrained upper confidence reinforcement learning
Liyuan Zheng and Lillian Ratliff · 2020
Cited alongside, same era.
A sample-efficient algorithm for episodic finite-horizon mdp with constraints
Krishna C Kalagarla, Rahul Jain, and Pierluigi Nuzzo · 2020
Cited alongside, same era.
Natural policy gradient primal-dual method for constrained markov decision processes
Dongsheng Ding, Kaiqing Zhang, Tamer Basar, and Mihailo R Jovanovic · 2020
Cited alongside, same era.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I Jordan · 2020
Cited alongside, same era.
Crpo: A new approach for safe reinforcement learning with convergence guarantee
Tengyu Xu, Yingbin Liang, and Guanghui Lan · 2021
Later among the works it cites.
Achieving zero constraint violation for constrained reinforcement learning via primal-dual approach
Qinbo Bai, Amrit Singh Bedi, Mridul Agarwal, Alec Koppel, and Vaneet Aggarwal · 2021
Later among the works it cites.
A provably-efficient model-free algorithm for constrained markov decision processes
Honghao Wei, Xin Liu, and Lei Ying · 2021
Later among the works it cites.
Safe reinforcement learning with linear function approximation
Sanae Amani, Christos Thrampoulidis, and Lin F Yang · 2021
Later among the works it cites.
Provably efficient reinforcement learning for discounted mdps with feature mapping
Dongruo Zhou, Jiafan He, and Quanquan Gu · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shuang Qiu, Xiaohan Wei, Zhuoran Yang, Jieping Ye, and Zhaoran Wang · 2020
Cited alongside, same era.
Model-based reinforcement learning with value-targeted regression
Alex Ayoub, Zeyu Jia, Csaba Szepesvari, Mengdi Wang, and Lin Yang · 2020
Cited alongside, same era.
Flambe: Structural complexity and representation learning of low rank mdps
Alekh Agarwal, Sham Kakade, Akshay Krishnamurthy, and Wen Sun · 2020
Cited alongside, same era.
Optimal approximation–smoothness tradeoffs for soft-max functions
Alessandro Epasto, Mohammad Mahdian, Vahab Mirrokni, and Manolis Zampetakis · 2020
Cited alongside, same era.
Qiaomin Xie, Yudong Chen, Zhaoran Wang, and Zhuoran Yang · 2020
Cited alongside, same era.
Learning policies with zero or bounded constraint violation for constrained mdps
Tao Liu, Ruida Zhou, Dileep Kalathil, PR Kumar, and Chao Tian · 2021
Cited alongside, same era.
Provably efficient safe exploration via primal-dual policy optimization
Dongsheng Ding, Xiaohan Wei, Zhuoran Yang, Zhaoran Wang, and Mihailo Jovanovic · 2021
Cited alongside, same era.
Later among the works it cites.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2021
Later among the works it cites.
Model-free representation learning and exploration in low-rank mdps
Aditya Modi, Jinglin Chen, Akshay Krishnamurthy, Nan Jiang, and Alekh Agarwal · 2021
Later among the works it cites.
A simple reward-free approach to constrained reinforcement learning
Sobhan Miryoosefi and Chi Jin · 2022
Closest in time.
Efficient reinforcement learning in block mdps: A model-free representation learning approach
Xuezhou Zhang, Yuda Song, Masatoshi Uehara, Mengdi Wang, Alekh Agarwal, and Wen Sun · 2022
Closest in time.
Near-optimal sample complexity bounds for constrained mdps
Sharan Vaswani, Lin F Yang, and Csaba Szepesvári · 2022
Closest in time.
Yuhao Ding and Javad Lavaei · 2022
Closest in time.