Fetching the paper…
Reading the bibliography…
We study policy optimization in an infinite horizon, $\gamma$-discounted constrained Markov decision process (CMDP).
The equivalence of two extremum problems
Jack Kiefer and Jacob Wolfowitz · 1960
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Constrained Markov decision processes , volume 7
Eitan Altman · 1999
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2001
Earlier work this paper cites.
Qos and fairness constrained convex optimization of resource allocation for wireless cellular and ad hoc networks
David Julian, Mung Chiang, Daniel O’Neill, and Stephen Boyd · 2002
Earlier work this paper cites.
Least-squares policy iteration
Michail G Lagoudakis and Ronald Parr · 2003
Earlier work this paper cites.
An actor-critic algorithm for constrained Markov decision processes
Vivek S Borkar · 2005
Earlier work this paper cites.
Modeling medical treatment using Markov decision processes
Andrew J Schaefer, Matthew D Bailey, Steven M Shechter, and Mark S Roberts · 2005
Earlier work this paper cites.
An overview on wireless sensor networks technology and evolution
Chiara Buratti, Andrea Conti, Davide Dardari, and Roberto Verdone · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
An online actor–critic algorithm with function approximation for constrained Markov decision processes
Shalabh Bhatnagar and K Lakshmanan · 2012
Earlier work this paper cites.
Safe exploration in Markov decision processes
Teodor Mihai Moldovan and Pieter Abbeel · 2012
Earlier work this paper cites.
Risk-constrained Markov decision processes
Vivek Borkar and Rahul Jain · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Chance-constrained dynamic programming with application to risk-aware robotic space exploration
Masahiro Ono, Marco Pavone, Yoshiaki Kuwata, and J Balaram · 2015
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Coin betting and parameter-free online learning
Francesco Orabona and David Pal · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Cited alongside, same era.
A theory of regularized Markov decision processes
Matthieu Geist, Bruno Scherrer, and Olivier Pietquin · 2019
Later among the works it cites.
A modern introduction to online learning
Francesco Orabona · 2019
Later among the works it cites.
Constrained reinforcement learning has zero duality gap
Santiago Paternain, Luiz FO Chamon, Miguel Calvo-Fullana, and Alejandro Ribeiro · 2019
Later among the works it cites.
Convergent policy optimization for safe reinforcement learning
Ming Yu, Zhuoran Yang, Mladen Kolar, and Zhaoran Wang · 2019
Later among the works it cites.
Natural policy gradient primal-dual method for constrained Markov decision processes
Dongsheng Ding, Kaiqing Zhang, Tamer Basar, and Mihailo Jovanovic · 2020
Later among the works it cites.
IPO: Interior-point policy optimization under constraints
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Francesco Orabona and Tatiana Tommasi · 2017
Cited alongside, same era.
A Lyapunov-based approach to safe reinforcement learning
Yinlam Chow, Ofir Nachum, Edgar Duenez-Guzman, and Mohammad Ghavamzadeh · 2018
Cited alongside, same era.
Safe exploration in continuous action spaces
Gal Dalal, Krishnamurthy Dvijotham, Matej Vecerik, Todd Hester, Cosmin Paduraru, and Yuval Tassa · 2018
Cited alongside, same era.
A general safety framework for learning-based control in uncertain robotic systems
Jaime F Fisac, Anayo K Akametalu, Melanie N Zeilinger, Shahab Kaynama, Jeremy Gillula, and Claire J Tomlin · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Accelerated primal-dual policy optimization for safe reinforcement learning
Qingkai Liang, Fanyu Que, and Eytan Modiano · 2018
Cited alongside, same era.
Reinforcement learning: An introduction. second
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Yongshuai Liu, Jiaxin Ding, and Xin Liu · 2020
Later among the works it cites.
Responsive safety in reinforcement learning by PID Lagrangian methods
Adam Stooke, Joshua Achiam, and Pieter Abbeel · 2020
Later among the works it cites.
Tossingbot: Learning to throw arbitrary objects with residual physics
Andy Zeng, Shuran Song, Johnny Lee, Alberto Rodriguez, and Thomas Funkhouser · 2020
Later among the works it cites.
Fast global convergence of natural policy gradient methods with entropy regularization
Shicong Cen, Chen Cheng, Yuxin Chen, Yuting Wei, and Yuejie Chi · 2021
Later among the works it cites.
Provably efficient safe exploration via primal-dual policy optimization
Dongsheng Ding, Xiaohan Wei, Zhuoran Yang, Zhaoran Wang, and Mihailo Jovanovic · 2021
Later among the works it cites.
Sample-efficient reinforcement learning is feasible for linearly realizable MDPs with limited revisiting
Gen Li, Yuxin Chen, Yuejie Chi, Yuantao Gu, and Yuting Wei · 2021
Later among the works it cites.
A parameter-free algorithm for convex-concave min-max problems
Mingrui Liu and Francesco Orabona · 2021
Later among the works it cites.
CRPO: A new approach for safe reinforcement learning with convergence guarantee
Tengyu Xu, Yingbin Liang, and Guanghui Lan · 2021
Later among the works it cites.
Reward is enough for convex MDPs
Tom Zahavy, Brendan O’Donoghue, Guillaume Desjardins, and Satinder Singh · 2021
Later among the works it cites.