Fetching the paper…
Reading the bibliography…
Though deep reinforcement learning (DRL) has obtained substantial success, it may encounter catastrophic failures due to the intrinsic uncertainty of both transition and observation.
Consideration of risk in reinforcement learning
Matthias Heger · 1994
Earlier work this paper cites.
Nonlinear programming
Dimitri P Bertsekas · 1997
Earlier work this paper cites.
Optimization of conditional value-at-risk
R Tyrrell Rockafellar, Stanislav Uryasev, et al · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
A comparison of var and cvar constraints on portfolio selection with the mean-variance model
Gordon J Alexander and Alexandre M Baptista · 2004
Earlier work this paper cites.
Robust control of markov decision processes with uncertain transition matrices
Arnab Nilim and Laurent El Ghaoui · 2005
Earlier work this paper cites.
Minimizing cvar and var for a portfolio of derivatives
Siddharth Alexander, Thomas F Coleman, and Yuying Li · 2006
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Scaling up robust mdps by reinforcement learning
Aviv Tamar, Huan Xu, and Shie Mannor · 2013
Earlier work this paper cites.
Robust markov decision processes
Wolfram Wiesemann, Daniel Kuhn, and Berç Rustem · 2013
Earlier work this paper cites.
Algorithms for cvar optimization in mdps
Yinlam Chow and Mohammad Ghavamzadeh · 2014
Cited alongside, same era.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy · 2014
Cited alongside, same era.
Policy gradients for cvar-constrained mdps
LA Prashanth · 2014
Cited alongside, same era.
Risk-sensitive and robust decision-making: a cvar optimization approach
Yinlam Chow, Aviv Tamar, Shie Mannor, and Marco Pavone · 2015
Cited alongside, same era.
A comprehensive survey on safe reinforcement learning
Javier Garcıa and Fernando Fernández · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Adversarial attacks on neural network policies
Sandy Huang, Nicolas Papernot, Ian Goodfellow, Yan Duan, and Pieter Abbeel · 2017
Later among the works it cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Spinning Up in Deep Reinforcement Learning
Joshua Achiam · 2018
Later among the works it cites.
Implicit quantile networks for distributional reinforcement learning
Will Dabney, Georg Ostrovski, David Silver, and Rémi Munos · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Learning to drive in a day
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin A. Riedmiller, Andreas Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Optimizing the cvar via sampling
Aviv Tamar, Yonatan Glassner, and Shie Mannor · 2015
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Vedavyas Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy P. Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Cited alongside, same era.
Risk-constrained reinforcement learning with percentile risk criteria
Yinlam Chow, Mohammad Ghavamzadeh, Lucas Janson, and Marco Pavone · 2017
Cited alongside, same era.
Alex Kendall, Jeffrey Hawke, David Janz, Przemyslaw Mazur, Daniele Reda, John-Mark Allen, Vinh-Dieu Lam, Alex Bewley, and Amar Shah · 2019
Later among the works it cites.
Dsac: Distributional soft actor critic for risk-sensitive reinforcement learning
Xiaoteng Ma, Li Xia, Zhengyuan Zhou, Jun Yang, and Qianchuan Zhao · 2020
Later among the works it cites.
Worst cases policy gradients
Yichuan Charlie Tang, Jian Zhang, and Ruslan Salakhutdinov · 2020
Later among the works it cites.
Robust deep reinforcement learning against adversarial perturbations on state observations
Huan Zhang, Hongge Chen, Chaowei Xiao, Bo Li, Mingyan Liu, Duane Boning, and Cho-Jui Hsieh · 2020
Later among the works it cites.
Risk-averse bayes-adaptive reinforcement learning
Marc Rigter, Bruno Lacerda, and Nick Hawes · 2021
Later among the works it cites.
Wcsac: Worst-case soft actor critic for safety-constrained reinforcement learning
Qisong Yang, Thiago D Simão, Simon H Tindemans, and Matthijs TJ Spaan · 2021
Later among the works it cites.