Fetching the paper…
Reading the bibliography…
In many sequential decision-making problems we may want to manage risk by minimizing some measure of variability in rewards in addition to maximizing a standard criterion.
Mutual fund performance
W. Sharpe · 1966
Earlier work this paper cites.
On asymptotic normality in stochastic approximation
V. Fabian · 1968
Earlier work this paper cites.
Risk sensitive Markov decision processes
R. Howard and J. Matheson · 1972
Earlier work this paper cites.
Convergence of a class of random search algorithms
V. Katkovnik and Y. Kulchitsky · 1972
Earlier work this paper cites.
Stochastic approximation methods for constrained and unconstrained systems
H. Kushner and D. Clark · 1978
Earlier work this paper cites.
Practical optimization
P. Gill, W. Murray, and M. Wright · 1981
Earlier work this paper cites.
The variance of discounted Markov decision processes
M. Sobel · 1982
Earlier work this paper cites.
Neuron-like elements that can solve difficult learning control problems
A. Barto, R. Sutton, and C. Anderson · 1983
Earlier work this paper cites.
Temporal credit assignment in reinforcement learning
R. Sutton · 1984
Earlier work this paper cites.
Algorithms and software tools for IC yield optimization based on fundamental fabrication parameters
M. A. Styblinski and L. J. Opalski · 1986
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
R. Sutton · 1988
Earlier work this paper cites.
Variance-penalized Markov decision processes
J. Filar, L. Kallenberg, and H. Lee · 1989
Earlier work this paper cites.
Convergent activation dynamics in continuous time networks
M. W. Hirsch · 1989
Earlier work this paper cites.
Multivariate stochastic approximation using a simultaneous perturbation gradient approximation
J. Spall · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. Williams · 1992
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
M. Puterman · 1994
Earlier work this paper cites.
Dynamic Programming and Optimal Control
D. Bertsekas · 1995
Earlier work this paper cites.
Percentile performance criteria for limiting average Markov decision processes
J. Filar, D. Krass, and K. Ross · 1995
Earlier work this paper cites.
Microeconomic theory
A. Mas-Colell, M. Whinston, and J. Green · 1995
Earlier work this paper cites.
Neuro-Dynamic Programming
D. Bertsekas and J. Tsitsiklis · 1996
Earlier work this paper cites.
Weighted means in stochastic approximation of minima
J. Dippon and J. Renz · 1997
Earlier work this paper cites.
A one-measurement form of simultaneous perturbation stochastic approximation
J. Spall · 1997
Earlier work this paper cites.
An analysis of temporal-difference learning with function approximation
John N Tsitsiklis and Benjamin Van Roy · 1997
Cited alongside, same era.
Simulated-Based Methods for Markov Decision Processes
P. Marbach · 1998
Cited alongside, same era.
Reinforcement learning: An introduction
R. Sutton and A. Barto · 1998
Cited alongside, same era.
Constrained Markov decision processes , volume 7
Eitan Altman · 1999
Cited alongside, same era.
Nonlinear programming
D. Bertsekas · 1999
Cited alongside, same era.
A Kiefer-Wolfowitz algorithm with randomized differences
H. Chen, T. Duncan, and B. Pasik-Duncan · 1999
Cited alongside, same era.
Policy gradient methods for reinforcement learning with function approximation
Natural actor-critic
J. Peters, S. Vijayakumar, and S. Schaal · 2005
Later among the works it cites.
Adaptive Newton-based multivariate smoothed functional algorithms for simulation optimization
S. Bhatnagar · 2007
Later among the works it cites.
Incremental natural actor-Critic algorithms
S. Bhatnagar, R. Sutton, M. Ghavamzadeh, and M. Lee · 2007
Later among the works it cites.
A learning algorithm for risk-sensitive cost
A. Basu, T. Bhattacharyya, and V. Borkar · 2008
Later among the works it cites.
Stochastic approximation: a dynamical systems viewpoint
V. Borkar · 2008
Later among the works it cites.
Natural actor-critic algorithms
S. Bhatnagar, R. Sutton, M. Ghavamzadeh, and M. Lee · 2009
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Richard S Sutton, David A McAllester, Satinder P Singh, Yishay Mansour, et al · 1999
Cited alongside, same era.
The ode method for convergence of stochastic approximation and reinforcement learning
Vivek S Borkar and Sean P Meyn · 2000
Cited alongside, same era.
Actor-Critic algorithms
V. Konda and J. Tsitsiklis · 2000
Cited alongside, same era.
Adaptive stochastic approximation by the simultaneous perturbation method
J. Spall · 2000
Cited alongside, same era.
Policy gradient methods for reinforcement learning with function approximation
R. Sutton, D. McAllester, S. Singh, and Y. Mansour · 2000
Cited alongside, same era.
Infinite-horizon policy-gradient estimation
J. Baxter and P. Bartlett · 2001
Cited alongside, same era.
An actor–critic algorithm with function approximation for discounted cost constrained Markov decision processes
S. Bhatnagar · 2010
Later among the works it cites.
Learning algorithms for risk-sensitive control
V. Borkar · 2010
Later among the works it cites.
Percentile optimization for Markov decision processes with parameter uncertainty
E. Delage and S. Mannor · 2010
Later among the works it cites.
Risk-averse dynamic programming for Markov decision processes
A. Ruszczyński · 2010
Later among the works it cites.
Stochastic approximation algorithms for constrained optimization via simulation
S. Bhatnagar, N. Hemachandra, and V. Mishra · 2011
Later among the works it cites.
Mean-variance optimization in Markov decision processes
S. Mannor and J. Tsitsiklis · 2011
Later among the works it cites.
Reinforcement Learning With Function Approximation for Traffic Signal Control
L.A. Prashanth and S. Bhatnagar · 2011
Later among the works it cites.
An online actor-critic algorithm with function approximation for constrained Markov decision processes
S. Bhatnagar and K. Lakshmanan · 2012
Later among the works it cites.
Threshold Tuning Using Stochastic Optimization for Graded Signal Control
L.A. Prashanth and S. Bhatnagar · 2012
Later among the works it cites.
Policy gradients with variance related risk criteria
A. Tamar, D. Di Castro, and S. Mannor · 2012
Later among the works it cites.
Distributionally robust Markov decision processes
H. Xu and S. Mannor · 2012
Later among the works it cites.
Stochastic Recursive Algorithms for Optimization , volume 434
S. Bhatnagar, H. Prasad, and L.A. Prashanth · 2013
Later among the works it cites.
Actor-critic algorithms for risk-sensitive MDPs
L.A. Prashanth and M. Ghavamzadeh · 2013
Later among the works it cites.
Risk-sensitive Markov control processes
Y. Shen, W. Stannat, and K. Obermayer · 2013
Later among the works it cites.
Variance adjusted actor-critic algorithms
A. Tamar and S. Mannor · 2013
Later among the works it cites.
Nathaniel Korda and L.A. Prashanth · 2014
Closest in time.