Fetching the paper…
Reading the bibliography…
Constrained Markov decision processes (CMDPs) model scenarios of sequential decision making with multiple objectives that are increasingly important in many applications.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R · 1933
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Lai, T. L. and Robbins, H · 1985
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L · 1994
Earlier work this paper cites.
Constrained Markov Decision Processes , volume 7
Altman, E · 1999
Earlier work this paper cites.
Online regret bounds for a new reinforcement learning algorithm
Auer, P. and Ortner, R · 2005
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P · 2010
Earlier work this paper cites.
An empirical evaluation of thompson sampling
Chapelle, O. and Li, L · 2011
Earlier work this paper cites.
Analysis of thompson sampling for the multi-armed bandit problem
Agrawal, S. and Goyal, N · 2012
Earlier work this paper cites.
Thompson sampling: An asymptotically optimal finite-time analysis
Kaufmann, E., Korda, N., and Munos, R · 2012
Earlier work this paper cites.
Thompson sampling for contextual bandits with linear payoffs
Agrawal, S. and Goyal, N · 2013
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Osband, I., Russo, D., and Van Roy, B · 2013
Earlier work this paper cites.
Thompson sampling for learning parameterized markov decision processes
Gopalan, A. and Mannor, S · 2015
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R · 2017
Cited alongside, same era.
Why is posterior sampling better than optimism for reinforcement learning?
Osband, I. and Van Roy, B · 2017
Cited alongside, same era.
Learning unknown markov decision processes: A thompson sampling approach
Ouyang, Y., Gagrani, M., Nayyar, A., and Jain, R · 2017
Cited alongside, same era.
Is Q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I · 2018
Cited alongside, same era.
Linear stochastic bandits under safety constraints
Amani, S., Alizadeh, M., and Thrampoulidis, C · 2019
Cited alongside, same era.
Constrained episodic reinforcement learning in concave-convex and knapsack settings
Brantley, K., Dudik, M., Lykouris, T., Miryoosefi, S., Simchowitz, M., Slivkins, A., and Sun, W · 2020
Model-free reinforcement learning in infinite-horizon average-reward markov decision processes
Wei, C.-Y., Jafarnia-Jahromi, M., Luo, H., Sharma, H., and Jain, R · 2020
Later among the works it cites.
Constrained upper confidence reinforcement learning
Zheng, L. and Ratliff, L · 2020
Later among the works it cites.
Provably efficient safe exploration via primal-dual policy optimization
Ding, D., Wei, X., Yang, Z., Wang, Z., and Jovanovic, M · 2021
Later among the works it cites.
Learning with safety constraints: Sample complexity of reinforcement learning for constrained mdps
HasanzadeZonuzy, A., Bura, A., Kalathil, D., and Shakkottai, S · 2021
Later among the works it cites.
A sample-efficient algorithm for episodic finite-horizon mdp with constraints
Kalagarla, K. C., Jain, R., and Nuzzo, P · 2021
Later among the works it cites.
Stochastic bandits with linear constraints
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Natural policy gradient primal-dual method for constrained markov decision processes
Ding, D., Zhang, K., Basar, T., and Jovanovic, M · 2020
Cited alongside, same era.
Exploration-exploitation in constrained mdps
Efroni, Y., Mannor, S., and Pirotta, M · 2020
Cited alongside, same era.
Safe linear stochastic bandits
Khezeli, K. and Bitar, E · 2020
Cited alongside, same era.
Upper confidence primal-dual reinforcement learning for cmdp with adversarial loss
Qiu, S., Wei, X., Yang, Z., Ye, J., and Wang, Z · 2020
Cited alongside, same era.
Learning in markov decision processes under constraints
Singh, R., Gupta, A., and Shroff, N. B · 2020
Cited alongside, same era.
Online learning for stochastic shortest path model via posterior sampling
Jafarnia-Jahromi, M., Chen, L., Jain, R., and Luo, H
Cited in the paper.
Pacchiano, A., Ghavamzadeh, M., Bartlett, P., and Jiang, H · 2021
Later among the works it cites.
Regret guarantees for model-based reinforcement learning with long-term average constraints
Agarwal, M., Bai, Q., and Aggarwal, V · 2022
Later among the works it cites.
Achieving zero constraint violation for constrained reinforcement learning via primal-dual approach
Bai, Q., Bedi, A. S., Agarwal, M., Koppel, A., and Aggarwal, V · 2022
Later among the works it cites.
Dope: Doubly optimistic and pessimistic exploration for safe reinforcement learning
Bura, A., Hasanzadezonuzy, A., Kalathil, D., Shakkottai, S., and Chamberland, J.-F · 2022
Later among the works it cites.
Learning infinite-horizon average-reward markov decision processes with constraints
Chen, L., Jain, R., and Luo, H · 2022
Later among the works it cites.
Triple-q: A model-free algorithm for constrained reinforcement learning with sublinear regret and zero constraint violation
Wei, H., Liu, X., and Ying, L · 2022
Later among the works it cites.