Fetching the paper…
Reading the bibliography…
In many real-world reinforcement learning (RL) problems, besides optimizing the main objective function, an agent must concurrently avoid violating a number of constraints.
Linear and nonlinear programming
D. Luenberger, Y. Ye, et al · 1984
Earlier work this paper cites.
Dynamic programming and optimal control
D. Bertsekas · 1995
Earlier work this paper cites.
Noninear systems
Hassan K Khalil · 1996
Earlier work this paper cites.
Constrained Markov decision processes with total cost criteria: Lagrangian approach and dual linear program
Eitan Altman · 1998
Earlier work this paper cites.
Multi-criteria reinforcement learning
Z. Gábor and Z. Kalmár · 1998
Earlier work this paper cites.
Constrained Markov decision processes
E. Altman · 1999
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
S. Kakade and J. Langford · 2002
Earlier work this paper cites.
Lyapunov design for safe reinforcement learning
T. Perkins and A. Barto · 2002
Earlier work this paper cites.
Convex optimization
S. Boyd and L. Vandenberghe · 2004
Earlier work this paper cites.
Risk-sensitive reinforcement learning applied to control under constraints
P. Geibel and F. Wysotzki · 2005
Earlier work this paper cites.
An MDP-based recommender system
G. Shani, D. Heckerman, and R. Brafman · 2005
Earlier work this paper cites.
On the complexity of learning lexicographic strategies
M. Schmitt and L. Martignon · 2006
Earlier work this paper cites.
Bounding stationary expectations of markov processes
P. Glynn, A. Zeevi, et al · 2008
Earlier work this paper cites.
Regret-based reward elicitation for Markov decision processes
K. Regan and C. Boutilier · 2009
Cited alongside, same era.
Optimizing debt collections using constrained reinforcement learning
N. Abe, P. Melville, C. Pendus, C. Reddy, D. Jensen, V. Thomas, J. Bennett, G. Anderson, B. Cooley, M. Kowalczyk, et al · 2010
Cited alongside, same era.
Stochastic network optimization with application to communication and queueing systems
M. Neely · 2010
Cited alongside, same era.
Double Q-learning
H. van Hasselt · 2010
Cited alongside, same era.
Fast reinforcement learning for energy-efficient wireless communication
N. Mastronarde and M. van der Schaar · 2011
Cited alongside, same era.
Safe exploration in Markov decision processes
T. Moldovan and P. Abbeel · 2012
A. Rusu, S. Colmenarejo, C. Gulcehre, G. Desjardins, J. Kirkpatrick, R. Pascanu, V. Mnih, K. Kavukcuoglu, and R. Hadsell · 2015
Later among the works it cites.
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2015
Later among the works it cites.
Concrete problems in AI safety
D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané · 2016
Later among the works it cites.
Convex synthesis of randomized policies for controlled Markov chains with density safety upper bound constraints
M El Chamie, Y. Yu, and B. Açıkmeşe · 2016
Later among the works it cites.
Continuous deep Q-learning with model-based acceleration
S. Gu, T. Lillicrap, I. Sutskever, and S. Levine · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Policy gradients with variance related risk criteria
A. Tamar, D. Di Castro, and S. Mannor · 2012
Cited alongside, same era.
Playing Atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Cited alongside, same era.
A survey of multi-objective sequential decision-making
D. Roijers, P. Vamplew, S. Whiteson, and R. Dazeley · 2013
Cited alongside, same era.
Performance bounds for λ \lambda policy iteration and application to the game of Tetris
Bruno Scherrer · 2013
Cited alongside, same era.
Trading safety versus performance: Rapid deployment of robotic swarms with robust performance constraints
Y. Chow, M. Pavone, B. Sadler, and S. Carpin · 2015
Cited alongside, same era.
Risk-constrained reinforcement learning with percentile risk criteria
Yinlam Chow, Mohammad Ghavamzadeh, Lucas Janson, and Marco Pavone · 2015
Cited alongside, same era.
Later among the works it cites.
Safety-constrained reinforcement learning for MDPs
S. Junges, N. Jansen, C. Dehnert, U. Topcu, and J. Katoen · 2016
Later among the works it cites.
Multi-objective deep reinforcement learning
H. Mossalam, Y. Assael, D. Roijers, and S. Whiteson · 2016
Later among the works it cites.
Constrained policy optimization
J. Achiam, D. Held, A. Tamar, and P. Abbeel · 2017
Later among the works it cites.
Safe model-based reinforcement learning with stability guarantees
F. Berkenkamp, M. Turchetta, A. Schoellig, and A. Krause · 2017
Later among the works it cites.
First-order methods almost always avoid saddle points
J. Lee, I. Panageas, G. Piliouras, M. Simchowitz, M. Jordan, and B. Recht · 2017
Later among the works it cites.
J. Leike, M. Martic, V. Krakovna, P. Ortega, T. Everitt, A. Lefrancq, L. Orseau, and S. Legg · 2017
Later among the works it cites.
A unified view of entropy-regularized markov decision processes
G. Neu, A. Jonsson, and V. Gómez · 2017
Later among the works it cites.
Safe exploration in continuous action spaces
G. Dalal, K. Dvijotham, M. Vecerik, T. Hester, C. Paduraru, and Y. Tassa · 2018
Closest in time.