Fetching the paper…
Reading the bibliography…
In safe reinforcement learning (SRL) problems, an agent explores the environment to maximize an expected total reward and meanwhile avoids violation of certain constraints on a number of expected total costs.
Ipo: interior-point policy optimization under constraints
Liu, Y., Ding, J., and Liu, X · 1910
Earlier work this paper cites.
Safe policies for reinforcement learning via primal-dual methods
Paternain, S., Calvo-Fullana, M., Chamon, L. F., and Ribeiro, A · 1911
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Constrained Markov decision processes , volume 7
Altman, E · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Earlier work this paper cites.
Qos and fairness constrained convex optimization of resource allocation for wireless cellular and ad hoc networks
Julian, D., Chiang, M., O’Neill, D., and Boyd, S · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J · 2002
Earlier work this paper cites.
A natural policy gradient
Kakade, S. M · 2002
Earlier work this paper cites.
Provably efficient safe exploration via primal-dual policy optimization
Ding, D., Wei, X., Yang, Z., Wang, Z., and Jovanović, M. R · 2003
Earlier work this paper cites.
Improving sample complexity bounds for actor-critic algorithms
Xu, T., Wang, Z., and Liang, Y · 2004
Earlier work this paper cites.
An actor-critic algorithm for constrained markov decision processes
Borkar, V. S · 2005
Earlier work this paper cites.
Non-asymptotic convergence analysis of two time-scale (natural) actor-critic algorithms
Xu, T., Wang, Z., and Liang, Y · 2005
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Nemirovski, A., Juditsky, A., Lan, G., and Shapiro, A · 2009
Earlier work this paper cites.
Weighted sums of random kitchen sinks: replacing minimization with randomization in learning
Rahimi, A. and Recht, B · 2009
Earlier work this paper cites.
An online actor–critic algorithm with function approximation for constrained markov decision processes
Bhatnagar, S. and Lakshmanan, K · 2012
Earlier work this paper cites.
An optimal method for stochastic composite optimization
Lan, G · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Ghadimi, S. and Lan, G · 2013
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Earlier work this paper cites.
OpenAI Gym, 2016
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Accelerated gradient methods for nonconvex nonlinear and stochastic programming
Ghadimi, S. and Lan, G · 2016
Earlier work this paper cites.
Algorithms for stochastic optimization with functional or expectation constraints
Lan, G. and Zhou, Z · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Levine, S., Finn, C., Darrell, T., and Abbeel, P · 2016
Cited alongside, same era.
Constrained policy optimization
Achiam, J., Held, D., Tamar, A., and Abbeel, P · 2017
Cited alongside, same era.
Risk-constrained reinforcement learning with percentile risk criteria
Chow, Y., Ghavamzadeh, M., Janson, L., and Pavone, M · 2017
Cited alongside, same era.
Adaptive batch size for safe policy gradients
Papini, M., Pirotta, M., and Restelli, M · 2017
Cited alongside, same era.
Optimality and approximation with policy gradient methods in Markov decision processes
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G · 2019
Later among the works it cites.
Global optimality guarantees for policy gradient methods
Bhandari, J. and Russo, D · 2019
Later among the works it cites.
Neural temporal-difference and q-learning provably converge to global optima
Cai, Q., Yang, Z., Lee, J. D., and Wang, Z · 2019
Later among the works it cites.
Lyapunov-based safe policy optimization for continuous control
Chow, Y., Nachum, O., Faust, A., Duenez-Guzman, E., and Ghavamzadeh, M · 2019
Later among the works it cites.
A tale of two-timescale reinforcement learning with the tightest finite-time bound
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Cited alongside, same era.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
Wu, Y., Mansimov, E., Grosse, R. B., Liao, S., and Ba, J · 2017
Cited alongside, same era.
A finite time analysis of temporal difference learning with linear function approximation
Bhandari, J., Russo, D., and Singal, R · 2018
Cited alongside, same era.
A Lyapunov-based approach to safe reinforcement learning
Chow, Y., Nachum, O., Duenez-Guzman, E., and Ghavamzadeh, M · 2018
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S., Zhai, X., Poczos, B., and Singh, A · 2018
Cited alongside, same era.
Global convergence of policy gradient methods for the linear quadratic regulator
Fazel, M., Ge, R., Kakade, S. M., and Mesbahi, M · 2018
Cited alongside, same era.
Dalal, G., Szorenyi, B., and Thoppe, G · 2019
Later among the works it cites.
Mastering the real-time strategy game starcraft ii
DeepMind, G. A · 2019
Later among the works it cites.
Kumar, H., Koppel, A., and Ribeiro, A · 2019
Later among the works it cites.
On the finite-time convergence of actor-critic algorithm
Qiu, S., Yang, Z., Ye, J., and Wang, Z · 2019
Later among the works it cites.
Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps
Shani, L., Efroni, Y., and Mannor, S · 2019
Later among the works it cites.
Hessian aided policy gradient
Shen, Z., Ribeiro, A., Hassani, H., Qian, H., and Mi, C · 2019
Later among the works it cites.
Finite-time error bounds for linear stochastic approximation and TD learning
Srikant, R. and Ying, L · 2019
Later among the works it cites.
Neural policy gradient methods: Global optimality and rates of convergence
Wang, L., Cai, Q., Yang, Z., and Wang, Z · 2019
Later among the works it cites.
Convergent policy optimization for safe reinforcement learning
Yu, M., Yang, Z., Kolar, M., and Wang, Z · 2019
Later among the works it cites.
Global convergence of policy gradient methods to (almost) locally optimal policies
Zhang, K., Koppel, A., Zhu, H., and Başar, T · 2019
Later among the works it cites.
A note on the linear convergence of policy gradient methods
Bhandari, J. and Russo, D · 2020
Closest in time.
Fast global convergence of natural policy gradient methods with entropy regularization
Cen, S., Cheng, C., Chen, Y., Wei, Y., and Chi, Y · 2020
Closest in time.
Single-timescale actor-critic provably finds globally optimal policy
Fu, Z., Yang, Z., and Wang, Z · 2020
Closest in time.
Responsive safety in reinforcement learning by pid lagrangian methods
Stooke, A., Achiam, J., and Abbeel, P · 2020
Closest in time.
Non-asymptotic convergence of adam-type reinforcement learning algorithms under Markovian sampling
Xiong, H., Xu, T., Liang, Y., and Zhang, W · 2020
Closest in time.
Zhang, Y., Cai, Q., Yang, Z., and Wang, Z · 2020
Closest in time.