Fetching the paper…
Reading the bibliography…
We consider a discounted cost constrained Markov decision process (CMDP) policy optimization problem, in which an agent seeks to maximize a discounted cumulative reward subject to a number of constraints on discounted cumulative utilities.
Finite-sample analysis for sarsa with linear function approximation
Zou, S., Xu, T., and Liang, Y. (2019) · 1902
Earlier work this paper cites.
On the global convergence of actor-critic: A case for linear quadratic regulator with ergodic cost
Yang, Z., Chen, Y., Hong, M., and Wang, Z. (2019) · 1907
Earlier work this paper cites.
Constrained reinforcement learning has zero duality gap
Paternain, S., Chamon, L. F., Calvo-Fullana, M., and Ribeiro, A. (2019) · 1910
Earlier work this paper cites.
Benchmarking safe exploration in deep reinforcement learning
Ray, A., Achiam, J., and Amodei, D. (2019) · 1910
Earlier work this paper cites.
Asymptotic properties of constrained markov decision processes
Altman, E. (1993) · 1993
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L. (1994) · 1994
Earlier work this paper cites.
Constrained Markov decision processes
Altman, E. (1999) · 1999
Earlier work this paper cites.
Finite-time analysis of stochastic gradient descent under markov randomness
Doan, T. T., Nguyen, L. M., Pham, N. H., and Romberg, J. (2020) · 2003
Earlier work this paper cites.
An actor-critic algorithm for constrained markov decision processes
Borkar, V. S. (2005) · 2005
Earlier work this paper cites.
A finite time analysis of two time-scale actor critic methods
Wu, Y., Zhang, W., Xu, P., and Gu, Q. (2020) · 2005
Earlier work this paper cites.
Markov chains and mixing times
Levin, D. A., Peres, Y., and Wilmer, E. L. (2006) · 2006
Earlier work this paper cites.
Dynamic programming in constrained markov decision processes
Piunovskiy, A. B. (2006) · 2006
Earlier work this paper cites.
Hong, M., Wai, H.-T., Wang, Z., and Yang, Z. (2020) · 2007
Earlier work this paper cites.
A natural actor-critic algorithm with downside risk constraints
Spooner, T. and Savani, R. (2020) · 2007
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey
Taylor, M. E. and Stone, P. (2009) · 2009
Earlier work this paper cites.
Optimizing debt collections using constrained reinforcement learning
Abe, N., Melville, P., Pendus, C., Reddy, C. K., Jensen, D. L., Thomas, V. P., Bennett, J. J., Anderson, G. F., Cooley, B. R., Kowalczyk, M., et al. (2010) · 2010
Earlier work this paper cites.
Zeng, S., Doan, T. T., and Romberg, J. (2020) · 2010
Earlier work this paper cites.
Nonlinear two-time-scale stochastic approximation: Convergence and finite-time performance
Doan, T. T. (2020) · 2011
Cited alongside, same era.
Neuroevolutionary reinforcement learning for generalized control of simulated helicopters
Koppejan, R. and Whiteson, S. (2011) · 2011
Cited alongside, same era.
Safe reinforcement learning for emergency loadshedding of power systems
Vu, T. L., Mukherjee, S., Yin, T., Huang, R., Huang, Q., et al. (2020) · 2011
Cited alongside, same era.
An online actor–critic algorithm with function approximation for constrained markov decision processes
Bhatnagar, S. and Lakshmanan, K. (2012) · 2012
Cited alongside, same era.
A novel q-learning algorithm with function approximation for constrained markov decision processes
Lakshmanan, K. and Bhatnagar, S. (2012) · 2012
Cited alongside, same era.
Finite-time performance bounds and adaptive learning rate selection for two time-scale reinforcement learning
Gupta, H., Srikant, R., and Ying, L. (2019) · 2019
Later among the works it cites.
Risk-aware motion planning and control using cvar-constrained optimization
Hakobyan, A., Kim, G. C., and Yang, I. (2019) · 2019
Later among the works it cites.
Finite-time error bounds for linear stochastic approximation andtd learning
Srikant, R. and Ying, L. (2019) · 2019
Later among the works it cites.
Convergent policy optimization for safe reinforcement learning
Yu, M., Yang, Z., Kolar, M., and Wang, Z. (2019) · 2019
Later among the works it cites.
Optimality and approximation with policy gradient methods in markov decision processes
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G. (2020) · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Risk-aware path planning using hirerachical constrained markov decision processes
Feyzabadi, S. and Carpin, S. (2014) · 2014
Cited alongside, same era.
A comprehensive survey on safe reinforcement learning
Garcıa, J. and Fernández, F. (2015) · 2015
Cited alongside, same era.
Chance-constrained dynamic programming with application to risk-aware robotic space exploration
Ono, M., Pavone, M., Kuwata, Y., and Balaram, J. (2015) · 2015
Cited alongside, same era.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P. (2017) · 2017
Cited alongside, same era.
Risk-constrained reinforcement learning with percentile risk criteria
Chow, Y., Ghavamzadeh, M., Janson, L., and Pavone, M. (2017) · 2017
Cited alongside, same era.
Robust adversarial reinforcement learning
Pinto, L., Davidson, J., Sukthankar, R., and Gupta, A. (2017) · 2017
Cited alongside, same era.
A lyapunov-based approach to safe reinforcement learning
Chow, Y., Nachum, O., Duenez-Guzman, E., and Ghavamzadeh, M. (2018) · 2018
Cited alongside, same era.
Anwar, A. and Raychowdhury, A. (2020) · 2020
Later among the works it cites.
Natural policy gradient primal-dual method for constrained markov decision processes
Ding, D., Zhang, K., Basar, T., and Jovanovic, M. R. (2020) · 2020
Later among the works it cites.
Teaching a humanoid robot to walk faster through safe reinforcement learning
García, J. and Shafie, D. (2020) · 2020
Later among the works it cites.
Safe learning and optimization techniques: Towards a survey of the state of the art
Kim, Y., Allmendinger, R., and López-Ibáñez, M. (2020) · 2020
Later among the works it cites.
On the global convergence rates of softmax policy gradient methods
Mei, J., Xiao, C., Szepesvari, C., and Schuurmans, D. (2020) · 2020
Later among the works it cites.
Improving generalization of reinforcement learning with minimax distributional soft actor-critic
Ren, Y., Duan, J., Li, S. E., Guan, Y., and Sun, Q. (2020) · 2020
Later among the works it cites.
Safe learning in robotics: From learning-based control to safe reinforcement learning
Brunke, L., Greeff, M., Hall, A. W., Yuan, Z., Zhou, S., Panerati, J., and Schoellig, A. P. (2021) · 2021
Closest in time.
Provably efficient safe exploration via primal-dual policy optimization
Ding, D., Wei, X., Yang, Z., Wang, Z., and Jovanovic, M. (2021) · 2021
Closest in time.
Finite sample analysis of two-time-scale natural actor-critic algorithm
Khodadadian, S., Doan, T. T., Maguluri, S. T., and Romberg, J. (2021) · 2021
Closest in time.
On finite-time convergence of actor-critic algorithm
Qiu, S., Yang, Z., Ye, J., and Wang, Z. (2021) · 2021
Closest in time.
Crpo: A new approach for safe reinforcement learning with convergence guarantee
Xu, T., Liang, Y., and Lan, G. (2021) · 2021
Closest in time.
Zeng, S., Doan, T. T., and Romberg, J. (2021) · 2021
Closest in time.