Fetching the paper…
Reading the bibliography…
We consider reinforcement learning with performance evaluated by a dynamic risk measure.
Simulation of self-organizing systems by digital computer
B. G. Farley and W. A. Clark · 1954
Earlier work this paper cites.
Theory of Neural-Analog Reinforcement Systems and Its Application to the Brain-Model Problem
M. L. Minsky · 1954
Earlier work this paper cites.
A Markovian decision process
R. E. Bellman · 1957
Earlier work this paper cites.
Dynamic Programming and Markov Processes
R. A. Howard · 1960
Earlier work this paper cites.
Polynomial approximation – a new computational technique in dynamic programming
Richard Bellman, Robert Kalaba, and Bella Kotkin · 1963
Earlier work this paper cites.
Risk-sensitive Markov decision processes
R. A. Howard and J. E. Matheson · 1971
Earlier work this paper cites.
Convergence conditions for nonlinear programming methods
E. A. Nurminski · 1972
Earlier work this paper cites.
Markov decision processes with a new optimality criterion: discrete time
S. C. Jaquette · 1973
Earlier work this paper cites.
Mean value theorem for convex functions
L. L. Wegge · 1974
Earlier work this paper cites.
A utility criterion for Markov decision processes
S. C. Jaquette · 1975
Earlier work this paper cites.
Optimal stopping, exponential utility, and linear programming
E. V. Denardo and U. G. Rothblum · 1979
Earlier work this paper cites.
Mean value theorems in nonsmooth analysis
JB Hiriart-Urruty · 1980
Earlier work this paper cites.
Stochastic approximation method with gradient averaging for unconstrained problems
A. Ruszczyński and W. Syski · 1983
Earlier work this paper cites.
Discounted MDPs: distribution functions and exponential utility maximization
K. J. Chung and M. J. Sobel · 1987
Earlier work this paper cites.
Learning to predict by the method of temporal differences
R. S. Sutton · 1988
Earlier work this paper cites.
Mean, variance, and probabilistic criteria in finite Markov decision processes: a review
D. J. White · 1988
Earlier work this paper cites.
Variance-penalized Markov decision processes
J. A. Filar, L. C. M. Kallenberg, and H.-M. Lee · 1989
Earlier work this paper cites.
Learning from Delayed Rewards
C. J. C. H. Watkins · 1989
Earlier work this paper cites.
The convergence of TD( λ \lambda ) for general λ \lambda
P. Dayan · 1992
Earlier work this paper cites.
Q - learning
C. J. C. H. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Efficient Dynamic Programming-Based Learning for Control
J. Peng · 1993
Earlier work this paper cites.
TD( λ \lambda ) converges with probability 1
P. Dayan and T. Sejnowski · 1994
Earlier work this paper cites.
On the convergence of stochastic iterative dynamic programming algorithms
T. Jaakkola, M. I. Jordan, and S. P. Singh · 1994
Earlier work this paper cites.
Incremental multi-step Q-learning
J. Peng and R. J. Williams · 1994
Earlier work this paper cites.
Markov Decision Processes
M. L. Puterman · 1994
Earlier work this paper cites.
On-line Q-learning using connectionist systems
G. A. Rummery and M. Niranjan · 1994
Earlier work this paper cites.
Asynchronous stochastic approximation and Q-learning
J. N. Tsitsiklis · 1994
Cited alongside, same era.
Problem Solving with Reinforcement Learning
G. A. Rummery · 1995
Cited alongside, same era.
Risk sensitive Markov decision processes
S. I. Marcus, E. Fernández-Gaucherand, D. Hernández-Hernández, S. Coraluppi, and P. Fard · 1997
Cited alongside, same era.
An analysis of temporal-difference learning with function approximation
J. N. Tsitsiklis and B. Van Roy · 1997
Cited alongside, same era.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 1998
Cited alongside, same era.
Coherent measures of risk
P. Artzner, F. Delbaen, J.-M. Eber, and D Heath · 1999
Cited alongside, same era.
Modeling, Measuring and Managing Risk
G.Ch. Pflug and W. Römisch · 2007
Later among the works it cites.
Valuations and dynamic convex risk measures
A. Jobert and L. C. G. Rogers · 2008
Later among the works it cites.
Risk-averse dynamic programming for Markov decision processes
A. Ruszczyński · 2010
Later among the works it cites.
Composition of time-consistent dynamic monetary risk measures in discrete time
P. Cheridito and M. Kupper · 2011
Later among the works it cites.
Approximate Dynamic Programming - Solving the Curses of Dimensionality
W. B. Powell · 2011
Later among the works it cites.
Policy gradients with variance related risk criteria
A. Tamar, D. Di Castro, and S. Mannor · 2012
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Risk sensitive control of finite state Markov chains in discrete time, with applications to portfolio management
T. Bielecki, D. Hernández-Hernández, and S. R. Pliska · 1999
Cited alongside, same era.
Risk-sensitive and minimax control of discrete-time, finite-state Markov decision processes
S. P. Coraluppi and S. I. Marcus · 1999
Cited alongside, same era.
Risk-sensitive control of discrete-time Markov processes with infinite horizon
G. B. Di Masi and Ł. Stettner · 1999
Cited alongside, same era.
Optimal long term growth rate of expected utility of wealth
W. H. Fleming and S. J. Sheu · 1999
Cited alongside, same era.
From stochastic dominance to mean–risk models: semideviations as risk measures
W. Ogryczak and A. Ruszczyński · 1999
Cited alongside, same era.
A sensitivity formula for risk-sensitive cost and the actor–critic algorithm
V.S. Borkar · 2001
Cited alongside, same era.
More risk-sensitive Markov decision processes
N. Bäuerle and U. Rieder · 2013
Later among the works it cites.
Persistently optimal policies in stochastic dynamic programming with generalized discounting
A. Jaśkiewicz, J. Matkowski, and A. S. Nowak · 2013
Later among the works it cites.
Dynamic programming with non-convex risk-sensitive measures
K. Lin and S. I. Marcus · 2013
Later among the works it cites.
Algorithmic aspects of mean-variance optimization in Markov decision processes
S. Mannor and J. N. Tsitsiklis · 2013
Later among the works it cites.
Risk-sensitive Markov control processes
Y. Shen, W. Stannat, and K. Obermayer · 2013
Later among the works it cites.
Markov decision problems where means bound variances
A. Arlotto, N. Gans, and J. M. Steele · 2014
Later among the works it cites.
Computational methods for risk-averse undiscounted transient Markov models
Ö. Çavus and A. Ruszczyński · 2014
Later among the works it cites.
Risk-averse control of undiscounted transient Markov models
Ö. Çavus and A. Ruszczyński · 2014
Later among the works it cites.
Time-consistent investment policies in Markovian markets: a case of mean-variance analysis
Z. Chen, G. Li, and Y. Zhao · 2014
Later among the works it cites.
Algorithms for CVaR optimization in MDPs
Y. Chow and M. Ghavamzadeh · 2014
Later among the works it cites.
Actor-critic algorithms for risk-sensitive reinforcement learning
L. A. Prashanth and Mohammad Ghavamzadeh · 2014
Later among the works it cites.
Scaling up robust mdps using function approximation
A. Tamar, S. Mannor, and H. Xu · 2014
Later among the works it cites.
Dynamic Programming and Optimal Control
D. P. Bersekas · 2017
Later among the works it cites.
Statistical estimation of composite risk functionals and risk optimization problems
D. Dentcheva, S. Penev, and A. Ruszczyński · 2017
Later among the works it cites.
Risk-averse sensor planning using distributed policy gradient
W.-J. Ma, D. Dentcheva, and M. M. Zavlanos · 2017
Later among the works it cites.
Sequential decision making with coherent risk
A. Tamar, Y. Chow, M. Ghavamzadeh, and S. Mannor · 2017
Later among the works it cites.
Process-based risk measures and risk-averse control of discrete-time systems
J. Fan and A. Ruszczyński · 2018
Later among the works it cites.
Risk measurement and risk-averse control of partially observable discrete-time markov systems
J. Fan and A. Ruszczyński · 2018
Later among the works it cites.
Risk forms: representation, disintegration, and application to partially observable two-stage systems
D. Dentcheva and A. Ruszczyński · 2019
Later among the works it cites.