Fetching the paper…
Reading the bibliography…
We study regret minimization for infinite-horizon average-reward Markov Decision Processes (MDPs) under cost constraints.
Markov decision processes: Discrete stochastic dynamic programming, 1994
Puterman, M. L · 1994
Earlier work this paper cites.
Constrained Markov decision processes , volume 7
Altman, E · 1999
Earlier work this paper cites.
Inequalities for the l1 deviation of the empirical distribution
Weissman, T., Ordentlich, E., Seroussi, G., Verdu, S., and Weinberger, M. J · 2003
Earlier work this paper cites.
Regal: A regularization based algorithm for reinforcement learning in weakly communicating mdps
Bartlett, P. L. and Tewari, A · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Jaksch, T., Ortner, R., and Auer, P · 2010
Earlier work this paper cites.
Conservative bandits
Wu, Y., Shariff, R., Lattimore, T., and Szepesvári, C · 2016
Earlier work this paper cites.
Conservative contextual linear bandits
Kazerouni, A., Ghavamzadeh, M., Abbasi-Yadkori, Y., and Van Roy, B · 2017
Earlier work this paper cites.
Accelerated primal-dual policy optimization for safe reinforcement learning
Liang, Q., Que, F., and Modiano, E · 2017
Earlier work this paper cites.
Efficient bias-span-constrained exploration-exploitation in reinforcement learning
Fruit, R., Pirotta, M., Lazaric, A., and Ortner, R · 2018
Earlier work this paper cites.
Is Q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., and Jordan, M. I · 2018
Earlier work this paper cites.
Regret bounds for reinforcement learning via Markov chain concentration
Ortner, R · 2018
Earlier work this paper cites.
Variance-aware regret bounds for undiscounted reinforcement learning in MDPs
Talebi, M. S. and Maillard, O.-A · 2018
Cited alongside, same era.
Reward constrained policy optimization
Tessler, C., Mankowitz, D. J., and Mannor, S · 2018
Cited alongside, same era.
Politex: Regret bounds for policy iteration using expert prediction
Abbasi-Yadkori, Y., Bartlett, P., Bhatia, K., Lazic, N., Szepesvari, C., and Weisz, G · 2019
Cited alongside, same era.
Linear stochastic bandits under safety constraints
Amani, S., Alizadeh, M., and Thrampoulidis, C · 2019
Cited alongside, same era.
Introduction to online convex optimization
Hazan, E · 2019
Cited alongside, same era.
Online convex optimization in adversarial Markov decision processes
Rosenberg, A. and Mansour, Y · 2019
Upper confidence primal-dual reinforcement learning for CMDP with adversarial loss
Qiu, S., Wei, X., Yang, Z., Ye, J., and Wang, Z · 2020
Later among the works it cites.
Optimistic policy optimization with bandit feedback
Shani, L., Efroni, Y., Rosenberg, A., and Mannor, S · 2020
Later among the works it cites.
Learning in Markov decision processes under constraints
Singh, R., Gupta, A., and Shroff, N. B · 2020
Later among the works it cites.
Model-free reinforcement learning in infinite-horizon average-reward Markov decision processes
Wei, C.-Y., Jahromi, M. J., Luo, H., Sharma, H., and Jain, R · 2020
Later among the works it cites.
Constrained upper confidence reinforcement learning
Zheng, L. and Ratliff, L · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Regret minimization for reinforcement learning by evaluating the optimal bias function
Zhang, Z. and Ji, X · 2019
Cited alongside, same era.
Near-optimal regret bounds for stochastic shortest path
Cohen, A., Kaplan, H., Mansour, Y., and Rosenberg, A · 2020
Cited alongside, same era.
Q-learning with UCB exploration is sample efficient for infinite-horizon MDP
Dong, K., Wang, Y., Chen, X., and Wang, L · 2020
Cited alongside, same era.
Exploration-exploitation in constrained mdps
Efroni, Y., Mannor, S., and Pirotta, M · 2020
Cited alongside, same era.
Improved algorithms for conservative exploration in bandits
Garcelon, E., Ghavamzadeh, M., Lazaric, A., and Pirotta, M · 2020
Cited alongside, same era.
Concave utility reinforcement learning with zero-constraint violations
Agarwal, M., Bai, Q., and Aggarwal, V
Cited in the paper.
Chen, L. and Luo, H · 2021
Later among the works it cites.
Minimax regret for stochastic shortest path
Cohen, A., Efroni, Y., Mansour, Y., and Rosenberg, A · 2021
Later among the works it cites.
A sample-efficient algorithm for episodic finite-horizon MDP with constraints
Kalagarla, K. C., Jain, R., and Nuzzo, P · 2021
Later among the works it cites.
Stochastic bandits with linear constraints
Pacchiano, A., Ghavamzadeh, M., Bartlett, P., and Jiang, H · 2021
Later among the works it cites.
Learning infinite-horizon average-reward MDPs with linear function approximation
Wei, C.-Y., Jahromi, M. J., Luo, H., and Jain, R · 2021
Later among the works it cites.