Fetching the paper…
Reading the bibliography…
We study the risk-sensitive exponential cost MDP formulation and develop a trajectory-based gradient algorithm to find the stationary point of the cost associated with a set of parameterized policies.
Optimization over time
Peter Whittle · 1982
Earlier work this paper cites.
Risk-sensitive optimal control
Peter Whittle · 1990
Earlier work this paper cites.
Connections between stochastic control and dynamic games
Paolo Dai Pra, Lorenzo Meneghini, and Wolfgang J Runggaldier · 1996
Earlier work this paper cites.
Multiplicative ergodicity and large deviations for an irreducible markov chain
S. Balaji and S.P. Meyn · 2000
Earlier work this paper cites.
A sensitivity formula for risk-sensitive cost and the actor–critic algorithm
V.S. Borkar · 2001
Earlier work this paper cites.
Simulation-based optimization of markov reward processes
P. Marbach and J.N. Tsitsiklis · 2001
Earlier work this paper cites.
Q-learning for risk-sensitive control
Vivek S Borkar · 2002
Earlier work this paper cites.
Risk-sensitive optimal control for markov decision processes with monotone cost
Vivek S Borkar and Sean P Meyn · 2002
Earlier work this paper cites.
Conditional value-at-risk for general loss distributions
R Tyrrell Rockafellar and Stanislav Uryasev · 2002
Earlier work this paper cites.
Spectral theory and limit theorems for geometrically ergodic Markov processes
I. Kontoyiannis and S. P. Meyn · 2003
Earlier work this paper cites.
A learning algorithm for risk-sensitive cost
Arnab Basu, Tirthankar Bhattacharyya, and Vivek S Borkar · 2008
Cited alongside, same era.
Stochastic Finance
Hans Föllmer and Alexander Schied · 2008
Cited alongside, same era.
Learning algorithms for risk-sensitive control
Vivek S Borkar · 2010
Cited alongside, same era.
Entropic risk measures: Coherence vs. convexity, model ambiguity and robust large deviations
Hans Föllmer and Thomas Knispel · 2011
Cited alongside, same era.
Robustness and risk-sensitivity in markov decision processes
Takayuki Osogami · 2012
Cited alongside, same era.
Actor-critic algorithms for risk-sensitive mdps
LA Prashanth and Mohammad Ghavamzadeh · 2013
Cited alongside, same era.
Variance-constrained actor-critic algorithms for discounted and average reward mdps
LA Prashanth and Mohammad Ghavamzadeh · 2016
Later among the works it cites.
Learning the variance of the reward-to-go
Aviv Tamar, Dotan Di Castro, and Shie Mannor · 2016
Later among the works it cites.
A variational formula for risk-sensitive reward
V. Anantharam and V. S. Borkar · 2017
Later among the works it cites.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Later among the works it cites.
Risk-constrained reinforcement learning with percentile risk criteria
Yinlam Chow, Mohammad Ghavamzadeh, Lucas Janson, and Marco Pavone · 2017
Later among the works it cites.
Risk-sensitive reinforcement learning: A constrained optimization viewpoint
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Temporal difference methods for the variance of the reward to go
Aviv Tamar, Dotan Di Castro, and Shie Mannor · 2013
Cited alongside, same era.
Algorithms for cvar optimization in mdps
Yinlam Chow and Mohammad Ghavamzadeh · 2014
Cited alongside, same era.
Risk-sensitive and robust decision-making: a cvar optimization approach
Yinlam Chow, Aviv Tamar, Shie Mannor, and Marco Pavone · 2015
Cited alongside, same era.
Optimizing the cvar via sampling
Aviv Tamar, Yonatan Glassner, and Shie Mannor · 2015
Cited alongside, same era.
Neuro-dynamic programming
Dimitri P Bertsekas and John N Tsitsiklis
Cited in the paper.
Michael Fu et al · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
Improving robustness via risk averse distributional reinforcement learning
Rahul Singh, Qinsheng Zhang, and Yongxin Chen · 2020
Later among the works it cites.
On tight bounds for function approximation error in risk-sensitive reinforcement learning
P. Karmakar and S. Bhatnagar · 2021
Later among the works it cites.
Kaiqing Zhang, Xiangyuan Zhang, Bin Hu, and Tamer Basar · 2021
Later among the works it cites.