Fetching the paper…
Reading the bibliography…
We study entropy-regularized constrained Markov decision processes (CMDPs) under the soft-max parameterization, in which an agent aims to maximize the entropy-regularized value function while satisfying constraints on the expected total utility.
Safe policies for reinforcement learning via primal-dual methods
Paternain, S., Calvo-Fullana, M., Chamon, L. F., and Ribeiro, A. (2019) · 1911
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K. (2016) · 1937
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate o (1/kˆ 2)
Nesterov, Y. E. (1983) · 1983
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Williams, R. J. and Peng, J. (1991) · 1991
Earlier work this paper cites.
Nonlinear and mixed-integer optimization: fundamentals and applications
Floudas, C. A. (1995) · 1995
Earlier work this paper cites.
Constrained Markov decision processes
Altman, E. (1999) · 1999
Earlier work this paper cites.
Elements of information theory
Cover, T. M. (1999) · 1999
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M. (2001) · 2001
Earlier work this paper cites.
Exploration-exploitation in constrained mdps
Efroni, Y., Mannor, S., and Pirotta, M. (2020) · 2003
Earlier work this paper cites.
An actor-critic algorithm for constrained markov decision processes
Borkar, V. S. (2005) · 2005
Earlier work this paper cites.
Constrained reinforcement learning from intrinsic and extrinsic rewards
Uchibe, E. and Doya, K. (2007) · 2007
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Ziebart, B. D. (2010) · 2010
Earlier work this paper cites.
Convergence rates of inexact proximal-gradient methods for convex optimization
Schmidt, M., Roux, N. L., and Bach, F. (2011) · 2011
Cited alongside, same era.
An online actor–critic algorithm with function approximation for constrained markov decision processes
Bhatnagar, S. and Lakshmanan, K. (2012) · 2012
Cited alongside, same era.
Perturbation analysis of optimization problems
Bonnans, J. F. and Shapiro, A. (2013) · 2013
Cited alongside, same era.
Gradient methods for minimizing composite functions
Nesterov, Y. (2013) · 2013
Cited alongside, same era.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P. (2017) · 2017
Cited alongside, same era.
Risk-constrained reinforcement learning with percentile risk criteria
Chow, Y., Ghavamzadeh, M., Janson, L., and Pavone, M. (2017) · 2017
Non-cooperative inverse reinforcement learning
Zhang, X., Zhang, K., Miehling, E., and Basar, T. (2019) · 2019
Later among the works it cites.
Natural policy gradient primal-dual method for constrained markov decision processes
Ding, D., Zhang, K., Basar, T., and Jovanovic, M. R. (2020) · 2020
Later among the works it cites.
On the global convergence rates of softmax policy gradient methods
Mei, J., Xiao, C., Szepesvari, C., and Schuurmans, D. (2020) · 2020
Later among the works it cites.
Teac: Intergrating trust region and max entropy actor critic for continuous control
Zang, H., Li, X., Zhang, L., Zhao, P., and Wang, M. (2020) · 2020
Later among the works it cites.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G. (2021) · 2021
Closest in time.
Fast global convergence of natural policy gradient methods with entropy regularization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S. (2017) · 2017
Cited alongside, same era.
Bridging the gap between value and policy based reinforcement learning
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D. (2017) · 2017
Cited alongside, same era.
A general safety framework for learning-based control in uncertain robotic systems
Fisac, J. F., Akametalu, A. K., Zeilinger, M. N., Kaynama, S., Gillula, J., and Tomlin, C. J. (2018) · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. (2018) · 2018
Cited alongside, same era.
Complexity of first-order inexact lagrangian and penalty methods for conic convex programming
Necoara, I., Patrascu, A., and Glineur, F. (2019) · 2019
Cited alongside, same era.
Cen, S., Cheng, C., Chen, Y., Wei, Y., and Chi, Y. (2021) · 2021
Closest in time.
A primal-dual approach to constrained markov decision processes
Chen, Y., Dong, J., and Wang, Z. (2021) · 2021
Closest in time.
Provably efficient safe exploration via primal-dual policy optimization
Ding, D., Wei, X., Yang, Z., Wang, Z., and Jovanovic, M. (2021) · 2021
Closest in time.
Faster algorithm and sharper analysis for constrained markov decision process
Li, T., Guan, Z., Zou, S., Xu, T., Liang, Y., and Lan, G. (2021) · 2021
Closest in time.
Leveraging non-uniformity in first-order non-convex optimization
Mei, J., Gao, Y., Dai, B., Szepesvari, C., and Schuurmans, D. (2021) · 2021
Closest in time.
Crpo: A new approach for safe reinforcement learning with convergence guarantee
Xu, T., Liang, Y., and Lan, G. (2021) · 2021
Closest in time.