Fetching the paper…
Reading the bibliography…
In many sequential decision-making problems we may want to manage risk by minimizing some measure of variability in costs in addition to minimizing a standard criterion.
Portfolio Selection: Efficient Diversification of Investment
H. Markowitz · 1959
Earlier work this paper cites.
Risk sensitive Markov decision processes
R. Howard and J. Matheson · 1972
Earlier work this paper cites.
The variance of discounted Markov decision processes
M. Sobel · 1982
Earlier work this paper cites.
Variance-penalized Markov decision processes
J. Filar, L. Kallenberg, and H. Lee · 1989
Earlier work this paper cites.
Multivariate stochastic approximation using a simultaneous perturbation gradient approximation
J. Spall · 1992
Earlier work this paper cites.
Percentile performance criteria for limiting average Markov decision processes
J. Filar, D. Krass, and K. Ross · 1995
Earlier work this paper cites.
Stochastic approximation algorithms and applications
Harold J Kushner and G George Yin · 1997
Earlier work this paper cites.
An integral invariance principle for differential inclusions with applications in adaptive control
EP Ryan · 1998
Earlier work this paper cites.
Coherent measures of risk
P. Artzner, F. Delbaen, J. Eber, and D. Heath · 1999
Earlier work this paper cites.
Nonlinear programming
D. Bertsekas · 1999
Earlier work this paper cites.
Conditional value-at-risk for general loss distributions
R. Rockafellar and S. Uryasev · 2000
Earlier work this paper cites.
A perturbation theory for ergodic markov chains and application to numerical approximations
Tony Shardlow and Andrew M Stuart · 2000
Earlier work this paper cites.
A sensitivity formula for the risk-sensitive cost and the actor-critic algorithm
V. Borkar · 2001
Cited alongside, same era.
Q-learning for risk-sensitive control
V. Borkar · 2002
Cited alongside, same era.
Nonlinear systems , volume 3
Hassan K Khalil and JW Grizzle · 2002
Cited alongside, same era.
Envelope theorems for arbitrary choice sets
Paul Milgrom and Ilya Segal · 2002
Cited alongside, same era.
Optimization of conditional value-at-risk
R. Rockafellar and S. Uryasev · 2002
Cited alongside, same era.
Perturbation analysis for denumerable markov chains with application to queueing models
Eitan Altman, Konstantin E Avrachenkov, and Rudesindo Núñez-Queija · 2004
Cited alongside, same era.
An actor-critic algorithm with function approximation for discounted cost constrained Markov decision processes
S. Bhatnagar · 2010
Later among the works it cites.
Nonparametric return distribution approximation for reinforcement learning
T. Morimura, M. Sugiyama, M. Kashima, H. Hachiya, and T. Tanaka · 2010
Later among the works it cites.
A Markov Decision Model for a Surveillance Application and Risk-Sensitive Markov Decision Processes
J. Ott · 2010
Later among the works it cites.
Markov decision processes with average-value-at-risk criteria
N. Bäuerle and J. Ott · 2011
Later among the works it cites.
An online actor-critic algorithm with function approximation for constrained Markov decision processes
S. Bhatnagar and K. Lakshmanan · 2012
Later among the works it cites.
An approximate solution method for large risk-averse Markov decision processes
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
An actor-critic algorithm for constrained Markov decision processes
V. Borkar · 2005
Cited alongside, same era.
Natural actor-critic
J. Peters, S. Vijayakumar, and S. Schaal · 2005
Cited alongside, same era.
Time consistent dynamic risk measures
K. Boda and J. Filar · 2006
Cited alongside, same era.
Stochastic approximation: a dynamical systems viewpoint
V. Borkar · 2008
Cited alongside, same era.
Computing VaR and CVaR using stochastic approximation and adaptive unconstrained importance sampling
O. Bardou, N. Frikha, and G. Pagès · 2009
Cited alongside, same era.
Natural actor-critic algorithms
S. Bhatnagar, R. Sutton, M. Ghavamzadeh, and M. Lee · 2009
Cited alongside, same era.
M. Petrik and D. Subramanian · 2012
Later among the works it cites.
Policy gradients with variance related risk criteria
A. Tamar, D. Di Castro, and S. Mannor · 2012
Later among the works it cites.
Stochastic Recursive Algorithms for Optimization , volume 434
S. Bhatnagar, H. Prasad, and L.A. Prashanth · 2013
Later among the works it cites.
Actor-critic algorithms for risk-sensitive MDPs
Prashanth L.A. and M. Ghavamzadeh · 2013
Later among the works it cites.
Risk-constrained Markov decision processes
V. Borkar and R. Jain · 2014
Closest in time.
Policy gradients beyond expectations: Conditional value-at-risk
A. Tamar, Y. Glassner, and S. Mannor · 2014
Closest in time.
Bias in natural actor-critic algorithms
P. Thomas · 2014
Closest in time.