Fetching the paper…
Reading the bibliography…
In order to model risk aversion in reinforcement learning, an emerging line of research adapts familiar algorithms to optimize coherent risk functionals, a class that includes conditional value-at-risk (CVaR).
Coherent measures of risk
Philippe Artzner, Freddy Delbaen, Jean-Marc Eber, and David Heath · 1999
Earlier work this paper cites.
Spectral measures of risk: A coherent representation of subjective risk aversion
Carlo Acerbi · 2002
Earlier work this paper cites.
Envelope theorems for arbitrary choice sets
Paul Milgrom and Ilya Segal · 2002
Earlier work this paper cites.
Minimizing cvar and var for a portfolio of derivatives
Siddharth Alexander, Thomas F Coleman, and Yuying Li · 2006
Earlier work this paper cites.
Optimization of convex risk functions
Andrzej Ruszczyński and Alexander Shapiro · 2006
Earlier work this paper cites.
Lectures on stochastic programming: Modeling and theory, mos-siam ser
A Shapiro, D Dentcheva, and A Ruszczynski · 2009
Earlier work this paper cites.
Risk-averse dynamic programming for markov decision processes
Andrzej Ruszczyński · 2010
Earlier work this paper cites.
Policy gradients with variance related risk criteria
Dotan Di Castro, Aviv Tamar, and Shie Mannor · 2012
Earlier work this paper cites.
Actor-critic algorithms for risk-sensitive mdps
LA Prashanth and Mohammad Ghavamzadeh · 2013
Earlier work this paper cites.
Algorithms for cvar optimization in mdps
Yinlam Chow and Mohammad Ghavamzadeh · 2014
Earlier work this paper cites.
Risk-sensitive and robust decision-making: a cvar optimization approach
Yinlam Chow, Aviv Tamar, Shie Mannor, and Marco Pavone · 2015
Cited alongside, same era.
A comprehensive survey on safe reinforcement learning
Javier Garcıa and Fernando Fernández · 2015
Cited alongside, same era.
CVXPY: A Python-embedded modeling language for convex optimization
Steven Diamond and Stephen Boyd · 2016
Cited alongside, same era.
Accelerated gradient methods for nonconvex nonlinear and stochastic programming
Saeed Ghadimi and Guanghui Lan · 2016
Cited alongside, same era.
Time-consistent decisions and temporal decomposition of coherent risk functionals
Georg Ch Pflug and Alois Pichler · 2016
Cited alongside, same era.
Cumulative prospect theory meets reinforcement learning: Prediction and control
LA Prashanth, Cheng Jie, Michael Fu, Steve Marcus, and Csaba Szepesvári · 2016
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Qiang Liu, Lihong Li, Ziyang Tang, and Dengyong Zhou · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift, 2019
Alekh Agarwal, Sham M. Kakade, Jason D. Lee, and Gaurav Mahajan · 2019
Later among the works it cites.
Global optimality guarantees for policy gradient methods
Jalaj Bhandari and Daniel Russo · 2019
Later among the works it cites.
Being optimistic to be conservative: Quickly learning a cvar policy
Ramtin Keramati, Christoph Dann, Alex Tamkin, and Emma Brunskill · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
First-order methods in optimization
Amir Beck · 2017
Cited alongside, same era.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Cited alongside, same era.
A rewriting system for convex optimization problems
Akshay Agrawal, Robin Verschueren, Steven Diamond, and Stephen Boyd · 2018
Cited alongside, same era.
Policy gradient in partially observable environments: Approximation and convergence
Kamyar Azizzadenesheli, Yisong Yue, and Animashree Anandkumar · 2018
Cited alongside, same era.
Policy gradient for coherent risk measures
Aviv Tamar, Yinlam Chow, Mohammad Ghavamzadeh, and Shie Mannor
Cited in the paper.
Optimizing the cvar via sampling
Aviv Tamar, Yonatan Glassner, and Shie Mannor
Cited in the paper.
Distributional reinforcement learning for efficient exploration
Borislav Mavrin, Shangtong Zhang, Hengshuai Yao, Linglong Kong, Kaiwen Wu, and Yaoliang Yu · 2019
Later among the works it cites.
Risk-sensitive reinforcement learning: Near-optimal risk-sample tradeoff in regret, 2020
Yingjie Fei, Zhuoran Yang, Yudong Chen, Zhaoran Wang, and Qiaomin Xie · 2020
Later among the works it cites.
Gradient descent-ascent provably converges to strict local minmax equilibria with a finite timescale separation, 2020
Tanner Fiez and Lillian Ratliff · 2020
Later among the works it cites.
Off-policy policy gradient with stationary distribution correction
Yao Liu, Adith Swaminathan, Alekh Agarwal, and Emma Brunskill · 2020
Later among the works it cites.