Fetching the paper…
Reading the bibliography…
Safety in reinforcement learning (RL) is a key property in both training and execution in many domains such as autonomous driving or finance.
Constrained Markov Decision Processes
E. Altman · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S. Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Convex Optimization
S. Boyd and L. Vandenberghe · 2004
Earlier work this paper cites.
Risk-sensitive reinforcement learning applied to control under constraints
P. Geibel and F. Wysotzky · 2005
Earlier work this paper cites.
Quantile Regression
R. Koenker · 2005
Earlier work this paper cites.
Learning algorithms for risk-sensitive control
V. S. Borkar · 2010
Earlier work this paper cites.
Safe policy iteration
M. Pirotta, M. Restelli, A. Pecorino, and D. Calandriello · 2013
Earlier work this paper cites.
Risk-constrained Markov decision processes
V. Borkar and R. Jain · 2014
Earlier work this paper cites.
Algorithms for CVaR optimization in MDPs
Y. Chow and M. Ghavamzadeh · 2014
Earlier work this paper cites.
Risk-Sensitive and Robust Decision-Making: a CVaR Optimization Approach
Y. Chow, A. Tamar, S. Mannor, and M. Pavone · 2015
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
J. Garcia and F. Fernandez · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, and D. Silver et al · 2015
Earlier work this paper cites.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M.I. Jordan, and P. Moritz · 2015
Earlier work this paper cites.
Variance-constrained actor-critic algorithms for discounted and average reward MDPs
L. Prashanth and M. Ghavamzadeh · 2016
Cited alongside, same era.
Safe exploration in finite Markov decision processes with gaussian processes
M. Turchetta, F. Berkenkamp, and A. Krause · 2016
Cited alongside, same era.
Constrained policy optimization
J. Achiam, D. Held, and A. Tamar et al · 2017
Cited alongside, same era.
A distributional perspective on reinforcement learning
M. Bellemare, W. Dabney, and R. Munos · 2017
Cited alongside, same era.
Safe model-based reinforcement learning with stability guarantees
F. Berkenkamp, M. Turchetta, A. Schoellig, and A. Krause · 2017
Cited alongside, same era.
Risk-constrained reinforcement learning with percentile risk criteria
Y. Chow, M. Ghavamzadeh, L. Janson, and M. Pavone · 2017
Cited alongside, same era.
Safe reinforcement learning via formal methods
N. Fulton and A. Platzer · 2018
Later among the works it cites.
Reinforcement learning: An introduction
R. Sutton and A. G. Barto · 2018
Later among the works it cites.
Safe exploration and optimization of constrained MDPs using gaussian processes
A. Wachi, Y. Sui, Y. Yue, and M. Ono · 2018
Later among the works it cites.
End-to-end safe reinforcement learning through barrier functions for safety-critical continuous control tasks
R. Cheng, G. Orosz, R. M Murray, and J. W Burdick · 2019
Later among the works it cites.
Reinforcement learning with convex constraints
S. Miryoosefi, K. Brantley, H. Daume III, M. Dudik, and R. Schapire · 2019
Later among the works it cites.
Benchmarking Safe Exploration in Deep Reinforcement Learning
A. Ray, J. Achiam, and D. Amodei · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep learning at Alibaba
R. Jin · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, and K. Simonyan et al · 2017
Cited alongside, same era.
Safe reinforcement learning via shielding
M. Alshiekh, R. Bloem, R. R. Ehlers, B. Könighofer, S. Niekum, and U. Topcu · 2018
Cited alongside, same era.
Distributed distributional deterministic policy gradients
G. Barth-Maron, M. W Hoffman, and D. et al. Budden · 2018
Cited alongside, same era.
Implicit quantile networks for distributional reinforcement learning
W. Dabney, G. Ostrovski, D. Silver, and R. Munos · 2018
Cited alongside, same era.
Reward constrained policy optimization
C. Tessler, D. Mankowitz, and S. Mannor · 2019
Later among the works it cites.
Fully parameterized quantile function for distributional reinforcement learning
D. Yang, L. Zhao, Z. Lin, T. Qin, J. Bian, and T. Liu · 2019
Later among the works it cites.
Convergent policy optimization for safe reinforcement learning
M. Yu, Z. Yang, M. Kolar, and Z. Wang · 2019
Later among the works it cites.
Reinforcement learning of risk-constrained policies in Markov decision processes
T. Brazdil, K. Chatterjee, P. Novotny, and J. Vahala · 2020
Later among the works it cites.
IPO: Interior-point policy optimization under constraints
Y. Liu, J. Ding, and X. Liu · 2020
Later among the works it cites.
Projection-based constrained policy optimization
T. Yang, J. Rosca, K. Narasimhan, and P. Ramadge · 2020
Later among the works it cites.