Fetching the paper…
Reading the bibliography…
In many sequential decision-making problems one is interested in minimizing an expected cumulative cost while taking into account \emph{risk}, i.e., increased awareness of events of small probability and high consequences.
Risk sensitive Markov decision processes
R. Howard and J. Matheson · 1972
Earlier work this paper cites.
The variance of discounted Markov decision processes
M. Sobel · 1982
Earlier work this paper cites.
Mean, variance, and probabilistic criteria in finite Markov decision processes: A review
D. White · 1988
Earlier work this paper cites.
Variance-penalized Markov decision processes
J. Filar, L. Kallenberg, and H. Lee · 1989
Earlier work this paper cites.
Multivariate stochastic approximation using a simultaneous perturbation gradient approximation
J. Spall · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. Williams · 1992
Earlier work this paper cites.
Dynamic programming and optimal control
D. Bertsekas · 1995
Earlier work this paper cites.
Percentile performance criteria for limiting average Markov decision processes
J. Filar, D. Krass, and K. Ross · 1995
Earlier work this paper cites.
Neuro-dynamic programming
D. Bertsekas and J. Tsitsiklis · 1996
Earlier work this paper cites.
Using Markov decision processes to optimize a nonlinear functional of the final distribution, with manufacturing applications
E. Collins · 1997
Earlier work this paper cites.
Stochastic approximation algorithms and applications
H. Kushner and G. Yin · 1997
Earlier work this paper cites.
Simulated-Based Methods for Markov Decision Processes
P. Marbach · 1998
Earlier work this paper cites.
Introduction to reinforcement learning
R. Sutton and A. Barto · 1998
Earlier work this paper cites.
Constrained Markov decision processes , volume 7
E. Altman · 1999
Earlier work this paper cites.
Coherent measures of risk
P. Artzner, F. Delbaen, J. Eber, and D. Heath · 1999
Earlier work this paper cites.
Nonlinear programming
D. Bertsekas · 1999
Earlier work this paper cites.
Minimizing risk models in Markov decision processes with policies depending on target values
C. Wu and Y. Lin · 1999
Earlier work this paper cites.
Actor-Critic algorithms
V. Konda and J. Tsitsiklis · 2000
Earlier work this paper cites.
Optimization of conditional value-at-risk
R. Rockafellar and S. Uryasev · 2000
Earlier work this paper cites.
A perturbation theory for ergodic Markov chains and application to numerical approximations
T. Shardlow and A. Stuart · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
R. Sutton, D. McAllester, S. Singh, and Y. Mansour · 2000
Cited alongside, same era.
Infinite-horizon policy-gradient estimation
J. Baxter and P. Bartlett · 2001
Cited alongside, same era.
A sensitivity formula for the risk-sensitive cost and the actor-critic algorithm
V. Borkar · 2001
Cited alongside, same era.
Q-learning for risk-sensitive control
V. Borkar · 2002
Cited alongside, same era.
Nonlinear systems , volume 3
H. Khalil and J. Grizzle · 2002
Cited alongside, same era.
Envelope theorems for arbitrary choice sets
P. Milgrom and I. Segal · 2002
Cited alongside, same era.
Natural actor-critic algorithms
S. Bhatnagar, R. Sutton, M. Ghavamzadeh, and M. Lee · 2009
Later among the works it cites.
An actor-critic algorithm with function approximation for discounted cost constrained Markov decision processes
S. Bhatnagar · 2010
Later among the works it cites.
Nonparametric return distribution approximation for reinforcement learning
T. Morimura, M. Sugiyama, M. Kashima, H. Hachiya, and T. Tanaka · 2010
Later among the works it cites.
A Markov decision model for a surveillance application and risk-sensitive Markov decision processes
J. Ott · 2010
Later among the works it cites.
Markov decision processes with average-value-at-risk criteria
N. Bäuerle and J. Ott · 2011
Later among the works it cites.
Value function approximation in reinforcement learning using the Fourier basis
G. Konidaris, S. Osentoski, and P. Thomas · 2011
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Rockafellar and S. Uryasev · 2002
Cited alongside, same era.
An MDP-based recommender system
G. Shani, R. Brafman, and D. Heckerman · 2002
Cited alongside, same era.
Perturbation analysis for denumerable Markov chains with application to queueing models
E. Altman, K. Avrachenkov, and R. Núñez-Queija · 2004
Cited alongside, same era.
Stochastic target hitting time and the problem of early retirement
K. Boda, J. Filar, Y. Lin, and L. Spanjers · 2004
Cited alongside, same era.
An actor-critic algorithm for constrained Markov decision processes
V. Borkar · 2005
Cited alongside, same era.
Natural actor-critic
J. Peters, S. Vijayakumar, and S. Schaal · 2005
Cited alongside, same era.
Later among the works it cites.
An online actor-critic algorithm with function approximation for constrained Markov decision processes
S. Bhatnagar and K. Lakshmanan · 2012
Later among the works it cites.
An approximate solution method for large risk-averse Markov decision processes
M. Petrik and D. Subramanian · 2012
Later among the works it cites.
Policy gradients with variance related risk criteria
A. Tamar, D. Di Castro, and S. Mannor · 2012
Later among the works it cites.
Stochastic recursive algorithms for optimization , volume 434
S. Bhatnagar, H. Prasad, and L. Prashanth · 2013
Later among the works it cites.
Stochastic Optimal Control with Dynamic, Time-Consistent Risk Constraints
Y. Chow and M. Pavone · 2013
Later among the works it cites.
Actor-critic algorithms for risk-sensitive MDPs
L. Prashanth and M. Ghavamzadeh · 2013
Later among the works it cites.
Risk neutral and risk averse stochastic dual dynamic programming method
A. Shapiro, W. Tekaya, J. da Costa, and M. Soares · 2013
Later among the works it cites.
Variance adjusted actor critic algorithms
A. Tamar and S. Mannor · 2013
Later among the works it cites.
Lifetime value marketing using reinforcement learning
G. Theocharous and A. Hallak · 2013
Later among the works it cites.
Risk-constrained Markov decision processes
V. Borkar and R. Jain · 2014
Later among the works it cites.
Algorithms for CVaR optimization in MDPs
Y. Chow and M. Ghavamzadeh · 2014
Later among the works it cites.
Chance-constrained dynamic programming with application to risk-aware robotic space exploration
M. Ono, M. Pavone, Y. Kuwata, and J. Balaram · 2015
Closest in time.
Policy gradients beyond expectations: Conditional value-at-risk
A. Tamar, Y. Glassner, and S. Mannor · 2015
Closest in time.