Fetching the paper…
Reading the bibliography…
Lagrangian methods are widely used algorithms for constrained optimization problems, but their learning dynamics exhibit oscillations and overshoot which, when applied to safe reinforcement learning, leads to constraint-violating behavior during agent training.
Lyapunov-based safe policy optimization for continuous control
Chow, Y., Nachum, O., Faust, A., Ghavamzadeh, M., and Duéñez-Guzmán, E. A · 1901
Earlier work this paper cites.
Multiplier and gradient methods
Hestenes, M. R · 1969
Earlier work this paper cites.
A method for nonlinear constraints in minimization problems
Powell, M. J · 1969
Earlier work this paper cites.
On penalty and multiplier methods for constrained minimization
Bertsekas, D. P · 1976
Earlier work this paper cites.
Constrained differential optimization
Platt, J. C. and Barr, A. H · 1988
Earlier work this paper cites.
Dynamic Systems Control: Linear Systems Analysis and Synthesis
Skelton, R · 1988
Earlier work this paper cites.
Nonlinear Control Systems
Isidori, A., Thoma, M., Sontag, E. D., Dickinson, B. W., Fettweis, A., Massey, J. L., and Modestino, J. W · 1995
Earlier work this paper cites.
Constrained markov decision processes with total cost criteria: Lagrangian approach and dual linear program
Altman, E · 1998
Earlier work this paper cites.
An optimal control model of neural networks for constrained optimization problems
Song, Q. and Leland, R. P · 1998
Earlier work this paper cites.
Introduction to Reinforcement Learning
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Constrained Markov decision processes , volume 7
Altman, E · 1999
Earlier work this paper cites.
How much uncertainty can be dealt with by feedback?
Liang-Liang Xie and Lei Guo · 2000
Earlier work this paper cites.
Improving the performance of weighted lagrange-multiplier methods for nonlinear constrained optimization
Wah, B. W., Wang, T., Shang, Y., and Wu, Z · 2000
Earlier work this paper cites.
Lipschitz continuity of optimal controls for state constrained problems
Galbraith, G. N. and Vinter, R. B · 2003
Cited alongside, same era.
Pid control
Åström, K. J. and Hägglund, T · 2006
Cited alongside, same era.
Numerical optimization
Nocedal, J. and Wright, S · 2006
Cited alongside, same era.
Risk-sensitive reinforcement learning applied to control under constraints
Geibel, P. and Wysotzki, F · 2011
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Cited alongside, same era.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Control interpretations for first-order optimization methods
Hu, B. and Lessard, L · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Later among the works it cites.
A pid controller approach for stochastic optimization of deep networks
An, W., Wang, H., Sun, Q., Xu, J., Dai, Q., and Zhang, L · 2018
Later among the works it cites.
A lyapunov-based approach to safe reinforcement learning
Chow, Y., Nachum, O., Duéñez-Guzmán, E. A., and Ghavamzadeh, M · 2018
Later among the works it cites.
Safe exploration in continuous action spaces
Dalal, G., Dvijotham, K., Vecerík, M., Hester, T., Paduraru, C., and Tassa, Y · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Constrained optimization and Lagrange multiplier methods
Bertsekas, D. P · 2014
Cited alongside, same era.
Analysis and design of optimization algorithms via integral quadratic constraints, 2014
Lessard, L., Recht, B., and Packard, A · 2014
Cited alongside, same era.
A general analysis of the convergence of admm, 2015
Nishihara, R., Lessard, L., Recht, B., Packard, A., and Jordan, M. I · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Cited alongside, same era.
Continuous deep q-learning with model-based acceleration
Gu, S., Lillicrap, T., Sutskever, I., and Levine, S · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Levine, S., Finn, C., Darrell, T., and Abbeel, P · 2016
Cited alongside, same era.
Later among the works it cites.
Openai five
OpenAI · 2018
Later among the works it cites.
Reward constrained policy optimization
Tessler, C., Mankowitz, D. J., and Mannor, S · 2018
Later among the works it cites.
Human-level performance in 3d multiplayer games with population-based reinforcement learning
Jaderberg, M., Czarnecki, W. M., Dunning, I., Marris, L., Lever, G., Castaneda, A. G., Beattie, C., Rabinowitz, N. C., Morcos, A. S., Ruderman, A., et al · 2019
Later among the works it cites.
Deep learning theory review: An optimal control and dynamical systems perspective, 2019
Liu, G.-H. and Theodorou, E. A · 2019
Later among the works it cites.
Ipo: Interior-point policy optimization under constraints
Liu, Y., Ding, J., and Liu, X · 2019
Later among the works it cites.
Constrained reinforcement learning has zero duality gap
Paternain, S., Chamon, L., Calvo-Fullana, M., and Ribeiro, A · 2019
Later among the works it cites.
Benchmarking Safe Exploration in Deep Reinforcement Learning
Ray, A., Achiam, J., and Amodei, D · 2019
Later among the works it cites.
Projection-based constrained policy optimization
Yang, T.-Y., Rosca, J., Narasimhan, K., and Ramadge, P. J · 2020
Closest in time.