Fetching the paper…
Reading the bibliography…
In contrast to the advances in characterizing the sample complexity for solving Markov decision processes (MDPs), the optimal statistical complexity for solving constrained MDPs (CMDPs) remains unknown.
Does knowledge transfer always help to learn a better policy?
Fei Feng, Wotao Yin, and Lin F. Yang · 1912
Earlier work this paper cites.
Constrained Markov decision processes , volume 7
Eitan Altman · 1999
Earlier work this paper cites.
Finite-sample convergence rates for q-learning and indirect algorithms
Michael Kearns and Satinder Singh · 1999
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham Machandranath Kakade · 2003
Earlier work this paper cites.
An actor-critic algorithm for constrained markov decision processes
Vivek S Borkar · 2005
Earlier work this paper cites.
Modeling medical treatment using markov decision processes
Andrew J Schaefer, Matthew D Bailey, Steven M Shechter, and Mark S Roberts · 2005
Earlier work this paper cites.
An overview on wireless sensor networks technology and evolution
Chiara Buratti, Andrea Conti, Davide Dardari, and Roberto Verdone · 2009
Earlier work this paper cites.
On the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar, Rémi Munos, and Bert Kappen · 2012
Earlier work this paper cites.
Minimax pac bounds on the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar, Rémi Munos, and Hilbert J Kappen · 2013
Earlier work this paper cites.
Risk-constrained markov decision processes
Vivek Borkar and Rahul Jain · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Earlier work this paper cites.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Cited alongside, same era.
Near-optimal time and sample complexities for solving markov decision processes with a generative model
Aaron Sidford, Mengdi Wang, Xian Wu, Lin F Yang, and Yinyu Ye · 2018
Cited alongside, same era.
Sim-to-real: Learning agile locomotion for quadruped robots
Jie Tan, Tingnan Zhang, Erwin Coumans, Atil Iscen, Yunfei Bai, Danijar Hafner, Steven Bohez, and Vincent Vanhoucke · 2018
Cited alongside, same era.
Reward constrained policy optimization
Chen Tessler, Daniel J Mankowitz, and Shie Mannor · 2018
Cited alongside, same era.
Constrained reinforcement learning has zero duality gap
Santiago Paternain, Luiz FO Chamon, Miguel Calvo-Fullana, and Alejandro Ribeiro · 2019
Achieving zero constraint violation for constrained reinforcement learning via primal-dual approach
Qinbo Bai, Amrit Singh Bedi, Mridul Agarwal, Alec Koppel, and Vaneet Aggarwal · 2021
Later among the works it cites.
A primal-dual approach to constrained markov decision processes
Yi Chen, Jing Dong, and Zhaoran Wang · 2021
Later among the works it cites.
Provably efficient safe exploration via primal-dual policy optimization
Dongsheng Ding, Xiaohan Wei, Zhuoran Yang, Zhaoran Wang, and Mihailo Jovanovic · 2021
Later among the works it cites.
Reinforcement learning for constrained markov decision processes
Ather Gattami, Qinbo Bai, and Vaneet Aggarwal · 2021
Later among the works it cites.
Model-based reinforcement learning for infinite-horizon discounted constrained markov decision processes
Aria HasanzadeZonuzy, Dileep M. Kalathil, and Srinivas Shakkottai · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Model-based reinforcement learning with a generative model is minimax optimal
Alekh Agarwal, Sham Kakade, and Lin F Yang · 2020
Cited alongside, same era.
Constrained episodic reinforcement learning in concave-convex and knapsack settings
Kianté Brantley, Miroslav Dudik, Thodoris Lykouris, Sobhan Miryoosefi, Max Simchowitz, Aleksandrs Slivkins, and Wen Sun · 2020
Cited alongside, same era.
Exploration-exploitation in constrained mdps
Yonathan Efroni, Shie Mannor, and Matteo Pirotta · 2020
Cited alongside, same era.
Breaking the sample size barrier in model-based reinforcement learning with a generative model
Gen Li, Yuting Wei, Yuejie Chi, Yuantao Gu, and Yuxin Chen · 2020
Cited alongside, same era.
Safe reinforcement learning in constrained markov decision processes
Akifumi Wachi and Yanan Sui · 2020
Cited alongside, same era.
Randomized linear programming solves the markov decision problem in nearly linear (sometimes sublinear) time
Mengdi Wang · 2020
Cited alongside, same era.
Tossingbot: Learning to throw arbitrary objects with residual physics
Andy Zeng, Shuran Song, Johnny Lee, Alberto Rodriguez, and Thomas Funkhouser · 2020
Cited alongside, same era.
Later among the works it cites.
A sample-efficient algorithm for episodic finite-horizon MDP with constraints
Krishna Chaitanya Kalagarla, Rahul Jain, and Pierluigi Nuzzo · 2021
Later among the works it cites.
A provably-efficient model-free algorithm for constrained markov decision processes
Honghao Wei, Xin Liu, and Lei Ying · 2021
Later among the works it cites.
On the sample complexity of batch reinforcement learning with policy-induced data
Chenjun Xiao, Ilbin Lee, Bo Dai, Dale Schuurmans, and Csaba Szepesvari · 2021
Later among the works it cites.
Crpo: A new approach for safe reinforcement learning with convergence guarantee
Tengyu Xu, Yingbin Liang, and Guanghui Lan · 2021
Later among the works it cites.
Provably efficient algorithms for multi-objective competitive rl
Tiancheng Yu, Yi Tian, Jingzhao Zhang, and Suvrit Sra · 2021
Later among the works it cites.
Towards painless policy optimization for constrained mdps
Arushi Jain, Sharan Vaswani, Reza Babanezhad, Csaba Szepesvari, and Doina Precup · 2022
Closest in time.
A simple reward-free approach to constrained reinforcement learning
Sobhan Miryoosefi and Chi Jin · 2022
Closest in time.