Fetching the paper…
Reading the bibliography…
Constrained Markov Decision Processes (CMDPs) formalize sequential decision-making problems whose objective is to minimize a cost function while satisfying constraints on various cost functions.
T. L. Lai and H. Robbins, “Asymptotically efficient adaptive allocation rules,” Advances in applied mathematics , vol. 6, no. 1, pp. 4–22, 1985
1985
Earlier work this paper cites.
M. L. Puterman, Markov Decision Processes: Discrete Stochastic Dynamic Programming , 1st ed. New York, NY, USA: John Wiley & Sons, Inc., 1994
1994
Earlier work this paper cites.
E. Altman, Constrained Markov Decision Processes . CRC Press, 1999, vol. 7
1999
Earlier work this paper cites.
R. I. Brafman and M. Tennenholtz, “R-max-a general polynomial time algorithm for near-optimal reinforcement learning,” Journal of Machine Learning Research , vol. 3, no. Oct, pp. 213–231, 2002
2002
Earlier work this paper cites.
V. S. Borkar, “An actor-critic algorithm for constrained markov decision processes,” Systems & control letters , vol. 54, no. 3, pp. 207–213, 2005
2005
Earlier work this paper cites.
P. Auer and R. Ortner, “Online regret bounds for a new reinforcement learning algorithm,” in Proceedings 1st Austrian Cognitive Vision Workshop , 2005
2005
Earlier work this paper cites.
P. Auer, T. Jaksch, and R. Ortner, “Near-optimal regret bounds for reinforcement learning,” in Advances in neural information processing systems , 2009, pp. 89–96
2009
Earlier work this paper cites.
A. L. Strehl, L. Li, and M. L. Littman, “Reinforcement learning in finite mdps: Pac analysis.” Journal of Machine Learning Research , vol. 10, no. 11, 2009
2009
Earlier work this paper cites.
2009
Earlier work this paper cites.
2009
Earlier work this paper cites.
A. Zimin and G. Neu, “Online learning in episodic markovian decision processes by relative entropy policy search,” in Advances in neural information processing systems , 2013, pp. 1583–1591
2013
Cited alongside, same era.
C. Dann and E. Brunskill, “Sample complexity of episodic fixed-horizon reinforcement learning,” in Advances in Neural Information Processing Systems , 2015, pp. 2818–2826
2015
Cited alongside, same era.
M. G. Azar, I. Osband, and R. Munos, “Minimax regret bounds for reinforcement learning,” in Proceedings of the 34th International Conference on Machine Learning - Volume 70 , ser. ICML’17, 2017, p. 263–272
2017
Cited alongside, same era.
C. Dann, T. Lattimore, and E. Brunskill, “Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning,” in Advances in Neural Information Processing Systems , 2017, pp. 5713–5723
2017
Cited alongside, same era.
S. Miryoosefi, K. Brantley, H. Daume III, M. Dudik, and R. E. Schapire, “Reinforcement learning with convex constraints,” in Advances in Neural Information Processing Systems , 2019, pp. 14 093–14 102
2019
Later among the works it cites.
2019
Later among the works it cites.
A. Rosenberg and Y. Mansour, “Online convex optimization in adversarial markov decision processes,” in International Conference on Machine Learning , 2019, pp. 5478–5486
2019
Later among the works it cites.
2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
C. Jin, Z. Allen-Zhu, S. Bubeck, and M. I. Jordan, “Is q-learning provably efficient?” in Advances in Neural Information Processing Systems , 2018, pp. 4863–4873
2018
Cited alongside, same era.
2018
Cited alongside, same era.
H. Le, C. Voloshin, and Y. Yue, “Batch policy learning under constraints,” ser. Proceedings of Machine Learning Research, vol. 97, 2019, pp. 3703–3712
2019
Cited alongside, same era.
A. Zanette and E. Brunskill, “Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds,” ser. Proceedings of Machine Learning Research, vol. 97. Long Beach, California, USA: PMLR, 09–15 Jun 2019, pp. 7304–7312
2019
Cited alongside, same era.
Y. Efroni, N. Merlis, M. Ghavamzadeh, and S. Mannor, “Tight regret bounds for model-based reinforcement learning with greedy policies,” in Advances in Neural Information Processing Systems , 2019, pp. 12 224–12 234
2019
Cited alongside, same era.
2020
Closest in time.
S. Qiu, X. Wei, Z. Yang, J. Ye, and Z. Wang, “Upper confidence primal-dual reinforcement learning for cmdp with adversarial loss,” 2020
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.