Fetching the paper…
Reading the bibliography…
Although Reinforcement Learning (RL) is effective for sequential decision-making problems under uncertainty, it still fails to thrive in real-world systems where risk or safety is a binding constraint.
Risk-sensitive Markov decision processes
Ronald A Howard and James E Matheson · 1972
Earlier work this paper cites.
On the Convergence Properties of the EM Algorithm
C. F. Jeff Wu · 1983
Earlier work this paper cites.
Asymptotic properties of a non-zero sum stochastic game
S Sorin · 1986
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
Ronald J Williams and Jing Peng · 1991
Earlier work this paper cites.
Consideration of risk in reinforcement learning
Matthias Heger · 1994
Earlier work this paper cites.
Risk sensitive markov decision processes
Steven I Marcus, Emmanual Fernández-Gaucherand, Daniel Hernández-Hernandez, Stefano Coraluppi, and Pedram Fard · 1997
Earlier work this paper cites.
Constrained Markov decision processes , volume 7
Eitan Altman · 1999
Earlier work this paper cites.
Coherent measures of risk
Philippe Artzner, Freddy Delbaen, Jean-Marc Eber, and David Heath · 1999
Earlier work this paper cites.
Risk-sensitive and minimax control of discrete-time, finite-state Markov decision processes
Stefano P Coraluppi and Steven I Marcus · 1999
Earlier work this paper cites.
Optimization of conditional value-at-risk
R Tyrrell Rockafellar, Stanislav Uryasev, et al · 2000
Earlier work this paper cites.
Risk measures for the 21st century , volume 1
Giorgio P Szegö · 2004
Earlier work this paper cites.
Risk analysis in robust control-making the case for probabilistic robust control
Xinjia Chen, Jorge L Aravena, and Kemin Zhou · 2005
Earlier work this paper cites.
Risk-sensitive reinforcement learning applied to control under constraints
Peter Geibel and Fritz Wysotzki · 2005
Earlier work this paper cites.
Relative entropy, exponential utility, and robust dynamic pricing
Andrew EB Lim and J George Shanthikumar · 2007
Earlier work this paper cites.
Linearly-solvable Markov decision problems
Emanuel Todorov · 2007
Earlier work this paper cites.
Robot trajectory optimization using approximate inference
Marc Toussaint · 2009
Earlier work this paper cites.
Double Q-learning
Hado Hasselt · 2010
Cited alongside, same era.
Algorithms for CVaR optimization in MDPs
Y. Chow and M. Ghavamzadeh · 2014
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Cited alongside, same era.
Risk-sensitive and robust decision-making: a cvar optimization approach
Y. Chow, A. Tamar, S. Mannor, and M. Pavone · 2015
Cited alongside, same era.
A comprehensive survey on safe reinforcement learning
Javier Garcıa and Fernando Fernández · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Later among the works it cites.
Soft actor-critic algorithms and applications
Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, George Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Learning-based model predictive control for safe exploration
Torsten Koller, Felix Berkenkamp, Matteo Turchetta, and Andreas Krause · 2018
Later among the works it cites.
Risk-sensitive reinforcement learning: A constrained optimization viewpoint
A Prashanth and Michael Fu · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Safe learning of regions of attraction for uncertain, nonlinear systems with Gaussian processes
Felix Berkenkamp, Riccardo Moriconi, Angela P. Schoellig, and Andreas Krause · 2016
Cited alongside, same era.
Google AI algorithm masters ancient game of Go
Elizabeth Gibney · 2016
Cited alongside, same era.
Variance-constrained actor-critic algorithms for discounted and average reward MDPs
LA Prashanth and Mohammad Ghavamzadeh · 2016
Cited alongside, same era.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Cited alongside, same era.
Risk-constrained reinforcement learning with percentile risk criteria
Yinlam Chow, Mohammad Ghavamzadeh, Lucas Janson, and Marco Pavone · 2017
Cited alongside, same era.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, et al · 2018
Later among the works it cites.
Understanding the impact of entropy on policy optimization
Zafarali Ahmed, Nicolas Le Roux, Mohammad Norouzi, and Dale Schuurmans · 2019
Later among the works it cites.
A differentiable physics engine for deep learning in robotics
Jonas Degrave, Michiel Hermans, Joni Dambre, et al · 2019
Later among the works it cites.
If MaxEnt RL is the answer, what is the question?
Benjamin Eysenbach and Sergey Levine · 2019
Later among the works it cites.
Benchmarking safe exploration in deep reinforcement learning
Alex Ray, Joshua Achiam, and Dario Amodei · 2019
Later among the works it cites.
An empirical investigation of the challenges of real-world reinforcement learning
Gabriel Dulac-Arnold, Nir Levine, Daniel J Mankowitz, Jerry Li, Cosmin Paduraru, Sven Gowal, and Todd Hester · 2020
Later among the works it cites.
Controlling overestimation bias with truncated mixture of continuous distributional quantile critics
Arsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, and Dmitry Vetrov · 2020
Later among the works it cites.
Safe reinforcement learning in constrained Markov decision processes
Akifumi Wachi and Yanan Sui · 2020
Later among the works it cites.
SENTINEL: Taming uncertainty with ensemble-based distributional reinforcement learning
Hannes Eriksson, Debabrota Basu, Mina Alibeigi, and Christos Dimitrakakis · 2021
Later among the works it cites.
Maximum entropy RL (provably) solves some robust RL problems
Benjamin Eysenbach and Sergey Levine · 2021
Later among the works it cites.
Recovery RL: Safe reinforcement learning with learned recovery zones
Brijen Thananjeyan, Ashwin Balakrishna, Suraj Nair, Michael Luo, Krishnan Srinivasan, Minho Hwang, Joseph E Gonzalez, Julian Ibarz, Chelsea Finn, and Ken Goldberg · 2021
Later among the works it cites.