Fetching the paper…
Reading the bibliography…
We study risk-sensitive RL where the goal is learn a history-dependent policy that optimizes some risk measure of cumulative rewards.
Portfolio selection, the journal of finance. 7 (1)
HM Markowitz · 1952
Earlier work this paper cites.
Risk-sensitive markov decision processes
Ronald A Howard and James E Matheson · 1972
Earlier work this paper cites.
The dual theory of choice under risk
Menahem E Yaari · 1987
Earlier work this paper cites.
Allais paradox
Maurice Allais · 1990
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
John Tsitsiklis and Benjamin Van Roy · 1996
Earlier work this paper cites.
Coherent measures of risk
Philippe Artzner, Freddy Delbaen, Jean-Marc Eber, and David Heath · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Optimization of conditional value-at-risk
R Tyrrell Rockafellar and Stanislav Uryasev · 2000
Earlier work this paper cites.
A natural policy gradient
Sham M Kakade · 2001
Earlier work this paper cites.
The reward hypothesis
Richard S. Sutton · 2004
Earlier work this paper cites.
An old-new concept of convex risk measures: The optimized certainty equivalent
Aharon Ben-Tal and Marc Teboulle · 2007
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
Risk-averse dynamic programming for markov decision processes
Andrzej Ruszczyński · 2010
Earlier work this paper cites.
Markov decision processes with average-value-at-risk criteria
Nicole Bäuerle and Jonathan Ott · 2011
Earlier work this paper cites.
Stochastic finance: an introduction in discrete time
Hans Föllmer and Alexander Schied · 2011
Earlier work this paper cites.
Prospect theory: An analysis of decision under risk
Daniel Kahneman and Amos Tversky · 2013
Earlier work this paper cites.
Robust markov decision processes
Wolfram Wiesemann, Daniel Kuhn, and Berç Rustem · 2013
Earlier work this paper cites.
Algorithms for cvar optimization in mdps
Yinlam Chow and Mohammad Ghavamzadeh · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Mimic-iii, a freely accessible critical care database
Alistair EW Johnson, Tom J Pollard, Lu Shen, Li-wei H Lehman, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark · 2016
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Risk-constrained reinforcement learning with percentile risk criteria
Yinlam Chow, Mohammad Ghavamzadeh, Lucas Janson, and Marco Pavone · 2018
Cited alongside, same era.
Implicit quantile networks for distributional reinforcement learning
Will Dabney, Georg Ostrovski, David Silver, and Rémi Munos · 2018
Cited alongside, same era.
On oracle-efficient pac rl with rich observations
Christoph Dann, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2018
Cited alongside, same era.
Provably filtering exogenous distractors using multistep inverse dynamics
Yonathan Efroni, Dipendra Misra, Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2022
Later among the works it cites.
Efficient risk-averse reinforcement learning
Ido Greenberg, Yinlam Chow, Mohammad Ghavamzadeh, and Shie Mannor · 2022
Later among the works it cites.
Mirror learning: A unifying framework of policy optimisation
Jakub Grudzien, Christian A Schroeder De Witt, and Jakob Foerster · 2022
Later among the works it cites.
Distributional reinforcement learning for risk-sensitive policies
Shiau Hong Lim and Ilyas Malik · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Representation learning for online and offline RL in low-rank MDPs
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nan Jiang and Alekh Agarwal · 2018
Cited alongside, same era.
Time limits in reinforcement learning
Fabio Pardo, Arash Tavakoli, Vitaly Levdik, and Petar Kormushev · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Andrea Zanette and Emma Brunskill · 2019
Cited alongside, same era.
Flambe: Structural complexity and representation learning of low rank mdps
Alekh Agarwal, Sham Kakade, Akshay Krishnamurthy, and Wen Sun · 2020
Cited alongside, same era.
Risk-sensitive reinforcement learning: Near-optimal risk-sample tradeoff in regret
Yingjie Fei, Zhuoran Yang, Yudong Chen, Zhaoran Wang, and Qiaomin Xie · 2020
Cited alongside, same era.
Revisiting design choices in proximal policy optimization
Chloe Ching-Yun Hsu, Celestine Mendler-Dünner, and Moritz Hardt · 2020
Cited alongside, same era.
Masatoshi Uehara, Xuezhou Zhang, and Wen Sun · 2022
Later among the works it cites.
On the convergence rates of policy gradient methods
Lin Xiao · 2022
Later among the works it cites.
Settling the reward hypothesis
Michael Bowling, John D Martin, David Abel, and Will Dabney · 2023
Later among the works it cites.
Provably efficient risk-sensitive reinforcement learning: Iterated CVar and worst path
Yihan Du, Siwei Wang, and Longbo Huang · 2023
Later among the works it cites.
Reinforcement learning in low-rank mdps with density features
Audrey Huang, Jinglin Chen, and Nan Jiang · 2023
Later among the works it cites.
Policy gradient for rectangular robust markov decision processes
Navdeep Kumar, Esther Derman, Matthieu Geist, Kfir Yehuda Levy, and Shie Mannor · 2023
Later among the works it cites.
Risk-aware reinforcement learning with coherent risk measures and non-linear function approximation
Thanh Lam, Arun Verma, Bryan Kian Hsiang Low, and Patrick Jaillet · 2023
Later among the works it cites.
One risk to rule them all: A risk-sensitive perspective on model-based offline reinforcement learning
Marc Rigter, Bruno Lacerda, and Nick Hawes · 2023
Later among the works it cites.
Near-minimax-optimal risk-sensitive reinforcement learning with cvar
Kaiwen Wang, Nathan Kallus, and Wen Sun · 2023
Later among the works it cites.
The role of coverage in online reinforcement learning
Tengyang Xie, Dylan J Foster, Yu Bai, Nan Jiang, and Sham M. Kakade · 2023
Later among the works it cites.
Regret bounds for markov decision processes with recursive optimized certainty equivalents
Wenhao Xu, Xuefeng Gao, and Xuedong He · 2023
Later among the works it cites.
Is risk-sensitive reinforcement learning properly resolved?
Ruiwen Zhou, Minghuan Liu, Kan Ren, Xufang Luo, Weinan Zhang, and Dongsheng Li · 2023
Later among the works it cites.
Efficient and sharp off-policy evaluation in robust markov decision processes
Andrew Bennett, Nathan Kallus, Miruna Oprescu, Wen Sun, and Kaiwen Wang · 2024
Closest in time.
The power of resets in online reinforcement learning
Zakaria Mhammedi, Dylan J Foster, and Alexander Rakhlin · 2024
Closest in time.
Position: A roadmap to pluralistic alignment
Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell L Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, et al · 2024
Closest in time.
Provably efficient CVar RL in low-rank MDPs
Yulai Zhao, Wenhao Zhan, Xiaoyan Hu, Ho fung Leung, Farzan Farnia, Wen Sun, and Jason D. Lee · 2024
Closest in time.