Fetching the paper…
Reading the bibliography…
In this paper we extend temporal difference policy evaluation algorithms to performance criteria that include the variance of the cumulative reward.
Mutual fund performance
Sharpe, W. F · 1966
Earlier work this paper cites.
Risk-sensitive markov decision processes
Howard, R. A. and Matheson, J. E · 1972
Earlier work this paper cites.
The variance of discounted markov decision processes
Sobel, M. J · 1982
Earlier work this paper cites.
Matrix Analysis
Horn, R. A. and Johnson, C. R · 1985
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 1994
Earlier work this paper cites.
Temporal difference learning and td-gammon
Tesauro, G · 1995
Earlier work this paper cites.
Neuro-dynamic programming
Bertsekas, D. P. and Tsitsiklis, J. N · 1996
Earlier work this paper cites.
Investment Science
Luenberger, D · 1998
Cited alongside, same era.
Reinforcement Learning
Sutton, R. S. and Barto, A. G · 1998
Cited alongside, same era.
Technical update: Least-squares temporal difference learning
Boyan, J.A · 2002
Cited alongside, same era.
Stochastic approximation: a dynamical systems viewpoint
Borkar, V.S · 2008
Cited alongside, same era.
Skill discovery in continuous reinforcement learning domains using skill chaining
Konidaris, G.D. and Barto, A.G · 2009
Cited alongside, same era.
Percentile optimization for Markov decision processes with parameter uncertainty
Delage, E. and Mannor, S · 2010
Cited alongside, same era.
Finite-sample analysis of lstd
Lazaric, A., Ghavamzadeh, M., and Munos, R · 2010
Later among the works it cites.
Temporal difference methods for general projected equations
Bertsekas, D.P · 2011
Later among the works it cites.
Mean-variance optimization in markov decision processes
Mannor, S. and Tsitsiklis, J. N · 2011
Later among the works it cites.
Dynamic Programming and Optimal Control, Vol II
Bertsekas, D. P · 2012
Later among the works it cites.
Parametric return density estimation for reinforcement learning
Morimura, T., Sugiyama, M., Kashima, H., Hachiya, H., and Tanaka, T · 2012
Later among the works it cites.
Policy gradients with variance related risk criteria
Tamar, A., Di Castro, D., and Mannor, S · 2012
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…