Fetching the paper…
Reading the bibliography…
The objective in a traditional reinforcement learning (RL) problem is to find a policy that optimizes the expected value of a performance metric such as the infinite-horizon cumulative discounted or long-run average cost/reward.
“Risk-sensitive control on an infinite time horizon”
W.. Fleming and W.. McEneaney · 1915
Earlier work this paper cites.
“Weighted bandits or: How bandits learn distorted values that are not expected”
A. Gopalan, L.. Prashanth, M.. Fu and S.. Marcus · 1947
Earlier work this paper cites.
“A stochastic approximation method”
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
“Stochastic estimation of the maximum of a regression function”
J. Kiefer and J. Wolfowitz · 1952
Earlier work this paper cites.
“Portfolio selection”
H. Markowitz · 1952
Earlier work this paper cites.
“Dynamic Programming”
Richard. Bellman · 1957
Earlier work this paper cites.
“Cost horizons and certainty equivalents: An approach to stochastic programming of heating oil”
A. Charnes, W.. Cooper and G.. Symonds · 1958
Earlier work this paper cites.
“On general minimax theorems”
M. Sion · 1958
Earlier work this paper cites.
“A Simulation-Based Algorithm for Ergodic Control of Markov Chains Conditioned on Rare Events”
S. Bhatnagar, V.. Borkar and M. Akarapu · 1962
Earlier work this paper cites.
“Chance-constrained programming with joint constraints”
L.. Miller and H. Wagner · 1965
Earlier work this paper cites.
“Stochastic optimization”
V.M. Aleksandrov, V.I. Sysoyev and V.V. Shemeneva · 1968
Earlier work this paper cites.
“Finite State Markovian Decision Processes”
Cyrus Derman · 1970
Earlier work this paper cites.
“On probabilistic constrained programming”
A. Prékopa · 1970
Earlier work this paper cites.
“Essays in the Theory of Risk Bearing”
K.. Arrow · 1971
Earlier work this paper cites.
“Risk-sensitive Markov decision processes”
R.. Howard and J.. Matheson · 1972
Earlier work this paper cites.
“Stochastic Approximation Methods for Constrained and Unconstrained Systems”
H. Kushner and D. Clark · 1978
Earlier work this paper cites.
“Prospect theory: An analysis of decision under risk”
D. Kahneman and A. Tversky · 1979
Earlier work this paper cites.
“The variance of discounted Markov decision processes”
M. Sobel · 1982
Earlier work this paper cites.
“Neuron-like elements that can solve difficult learning control problems”
A. Barto, R.. Sutton and C. Anderson · 1983
Earlier work this paper cites.
“Variance-reduced methods for machine learning”
R.. Gower, M. Schmidt, F. Bach and P. Richtárik · 1983
Earlier work this paper cites.
“Introduction to Stochastic Dynamic Programming”
S.. Ross · 1983
Earlier work this paper cites.
“Temporal Credit Assignment in Reinforcement Learning”, 1984
R.. Sutton · 1984
Earlier work this paper cites.
“Likelilood ratio gradient estimation: an overview”
Peter Glynn · 1987
Earlier work this paper cites.
“Learning to predict by the methods of temporal differences”
R.. Sutton · 1988
Earlier work this paper cites.
“An experimental test of several generalized utility theories”
Colin Camerer · 1989
Earlier work this paper cites.
“Three variants on the Allais example”
John Conlisk · 1989
Earlier work this paper cites.
“Variance-penalized Markov decision processes”
J. Filar, L. Kallenberg and H. Lee · 1989
Earlier work this paper cites.
“Sampling derivatives of probabilities”
G.. Pflug · 1989
Earlier work this paper cites.
“Sensitivity analysis for simulations via likelihood ratios”
M.I. Reiman and A. Weiss · 1989
Earlier work this paper cites.
“Sensitivity analysis of computer simulation models via the score efficient”
R.. Rubinstein · 1989
Earlier work this paper cites.
“Nonconvergence to unstable points in urn models and stochastic approximations”
R. Pemantle · 1990
Earlier work this paper cites.
“Risk-sensitive Optimal Control”, Wiley-Interscience series in systems and optimization
P. Whittle · 1990
Earlier work this paper cites.
“Stochastic approximation”
D. Ruppert · 1991
Earlier work this paper cites.
“Recent tests of generalizations of expected utility theory”
Colin Camerer · 1992
Earlier work this paper cites.
“Predictions about indifference curves inside the unit triangle: A test of variants of expected utility theory”
David Harless · 1992
Earlier work this paper cites.
“Acceleration of stochastic approximation by averaging”
B.. Polyak and A.. Juditsky · 1992
Earlier work this paper cites.
“Multivariate stochastic approximation using simultaneous perturbation gradient approximation”
J.. Spall · 1992
Earlier work this paper cites.
“Advances in prospect theory: Cumulative representation of uncertainty”
A. Tversky and D. Kahneman · 1992
Earlier work this paper cites.
“Discrete-time controlled Markov processes with average cost criterion: A survey”
A. Arapostathis et al · 1993
Earlier work this paper cites.
“A test of generalized expected utility theory”
Barry Sopher and Gary Gigliotti · 1993
Earlier work this paper cites.
“Violations of the betweenness axiom and nonlinearity in probability”
Colin Camerer and Teck-Hua Ho · 1994
Earlier work this paper cites.
“Markov Decision Processes: Discrete Stochastic Dynamic Programming”
M. Puterman · 1994
Earlier work this paper cites.
“Optimal Investment Policies for a Firm With a Random Risk Process: Exponential Utility and Minimizing the Probability of Ruin”
Sid Browne · 1995
Earlier work this paper cites.
“Percentile performance criteria for limiting average Markov decision processes”
J. Filar, D. Krass and K. Ross · 1995
Earlier work this paper cites.
“Microeconomic theory”
A. Mas-Colell, M. Whinston and J. Green · 1995
Earlier work this paper cites.
“Neuro-Dynamic Programming”
D.. Bertsekas and J.. Tsitsiklis · 1996
Earlier work this paper cites.
“Les algorithmes stochastiques contournent-ils les pieges?”
O. Brandiere and M. Duflo · 1996
Earlier work this paper cites.
“Risk sensitive control of Markov processes in countable state space”
D. Hernández-Hernández and S.. Marcus · 1996
Earlier work this paper cites.
“Optimization of Stochastic Models”
G.. Pflug · 1996
Earlier work this paper cites.
“Stochastic approximation with two time scales”
Vivek Borkar · 1997
Earlier work this paper cites.
“Risk-sensitive optimal control of hidden Markov models: Structural results”
E. Fernández-Gaucherand and S.. Marcus · 1997
Cited alongside, same era.
“Risk sensitive Markov decision processes”
S.. Marcus, E. Fernández-Gaucherand, S. D.ández-Hernández and P. Fard · 1997
Cited alongside, same era.
“An analysis of temporal-difference learning with function approximation”
J.. Tsitsiklis and B. Van · 1997
Cited alongside, same era.
“The probability weighting function”
Drazen Prelec · 1998
Cited alongside, same era.
“Constrained Markov Decision Processes”
E. Altman · 1999
Cited alongside, same era.
“Coherent measures of risk”
P. Artzner, F. Delbaen, J. Eber and D. Heath · 1999
Cited alongside, same era.
“An actor–critic algorithm with function approximation for discounted cost constrained Markov decision processes”
S. Bhatnagar · 2010
Later among the works it cites.
“Learning algorithms for risk-sensitive control”
V.. Borkar · 2010
Later among the works it cites.
“Risk-constrained Markov decision processes”
V.. Borkar and R. Jain · 2010
Later among the works it cites.
“Risk-averse dynamic programming for Markov decision processes”
A. Ruszczyński · 2010
Later among the works it cites.
“Deviation inequalities for an estimator of the conditional value-at-risk”
Y. Wang and F. Gao · 2010
Later among the works it cites.
“Infinite-horizon policy-gradient estimation”
Peter Bartlett and Jonathan Baxter · 2011
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Risk-Sensitive and Minimax Control of Discrete-Time, Finite-State Markov Decision Processes”
S.. Coraluppi and S.. Marcus · 1999
Cited alongside, same era.
“Risk-Sensitive, Minimax, and Mixed Risk Neutral/Minimax Control of Markov Decision Processes”
S.. Coraluppi and S.. Marcus · 1999
Cited alongside, same era.
“On the shape of the probability weighting function”
Richard Gonzalez and George Wu · 1999
Cited alongside, same era.
“Existence of risk sensitive optimal stationary policies for controlled Markov processes”
D. Hernández-Hernández and S.. Marcus · 1999
Cited alongside, same era.
“Actor-Critic–Type Learning Algorithms for Markov Decision Processes”
V.. Konda and V.. Borkar · 1999
Cited alongside, same era.
“Policy gradient methods for reinforcement learning with function approximation.”
R.. Sutton, D.. McAllester, S.. Singh and Y. Mansour · 1999
Cited alongside, same era.
Later among the works it cites.
“Reinforcement learning algorithms for MDPs”
C. Szepesvári · 2011
Later among the works it cites.
“Dynamic Programming and Optimal Control, Vol. II, 4th edition”
D.. Bertsekas · 2012
Later among the works it cites.
“Policy gradients with variance related risk criteria”
A. Tamar, D.. Castro and S. Mannor · 2012
Later among the works it cites.
“Thirty years of prospect theory in economics: A review and assessment”
N.. Barberis · 2013
Later among the works it cites.
“Stochastic Recursive Algorithms for Optimization”
S. Bhatnagar, H.. Prasad and L. A. Prashanth · 2013
Later among the works it cites.
“Simulation-based Algorithms for Markov Decision Processes”
H.. Chang, J. Hu, M.. Fu and S.. Marcus · 2013
Later among the works it cites.
“Stochastic first-and zeroth-order methods for nonconvex stochastic programming”
S. Ghadimi and G. Lan · 2013
Later among the works it cites.
“Stochastic Systems with Cumulative Prospect Theory”, 2013
K. Lin · 2013
Later among the works it cites.
“Cumulative weighting optimization: The discrete case”
K. Lin and S.. Marcus · 2013
Later among the works it cites.
“Dynamic programming with non-convex risk-sensitive measures”
K. Lin and S.. Marcus · 2013
Later among the works it cites.
“Algorithmic aspects of mean–variance optimization in Markov decision processes”
S. Mannor and J.. Tsitsiklis · 2013
Later among the works it cites.
“Actor-critic algorithms for risk-sensitive MDPs”
L.. Prashanth and M. Ghavamzadeh · 2013
Later among the works it cites.
“Risk-sensitive Markov control processes”
Y. Shen, W. Stannat and K. Obermayer · 2013
Later among the works it cites.
“Temporal difference methods for the variance of the reward to go”
A. Tamar, D. Di Castro and S. Mannor · 2013
Later among the works it cites.
“More risk-sensitive Markov decision processes”
N. Bäuerle and U. Rieder · 2014
Later among the works it cites.
“Risk-averse control of undiscounted transient Markov models”
O. Cavus and A. Ruszczynski · 2014
Later among the works it cites.
“Policy gradients for CVaR-constrained MDPs”
L.. Prashanth · 2014
Later among the works it cites.
“Lectures on Stochastic Programming: Modeling and Theory”
A. Shapiro, D. Dentcheva and A. Ruszczyński · 2014
Later among the works it cites.
“Optimizing the CVaR via sampling”
A. Tamar, Y. Glassner and S. Mannor · 2014
Later among the works it cites.
“Policy gradients beyond expectations: Conditional Value-at-Risk”
A. Tamar, Y. Glassner and S. Mannor · 2014
Later among the works it cites.
“Stochastic Gradient Estimation”
M.. Fu · 2015
Later among the works it cites.
“Escaping from saddle points—online stochastic gradient for tensor decomposition”
R. Ge, F. Huang, C. Jin and Y. Yuan · 2015
Later among the works it cites.
“Policy gradient for coherent risk measures”
A. Tamar, Y. Chow, M. Ghavamzadeh and S. Mannor · 2015
Later among the works it cites.
“Policy gradient for coherent risk measures”
A. Tamar, Y. Chow, M. Ghavamzadeh and S. Mannor · 2015
Later among the works it cites.
“Stochastic finance”
Hans Föllmer and Alexander Schied · 2016
Later among the works it cites.
“Variance-constrained actor-critic algorithms for discounted and average reward MDPs”
L.. Prashanth and M. Ghavamzadeh · 2016
Later among the works it cites.
“Cumulative prospect theory meets reinforcement learning: prediction and control”
L.. Prashanth et al · 2016
Later among the works it cites.
“Risk-constrained reinforcement learning with percentile risk criteria”
Y. Chow, M. Ghavamzadeh, L. Janson and M. Pavone · 2017
Later among the works it cites.
“Risk-averse approximate dynamic programming with quantile-based risk measures”
D.. Jiang and W.. Powell · 2017
Later among the works it cites.
“How to escape saddle points efficiently”
C. Jin et al · 2017
Later among the works it cites.
“Optimization methods for large-scale machine learning”
L. Bottou, F. Curtis and J. Nocedal · 2018
Closest in time.
“Stochastic optimization in a cumulative prospect theory framework”
C. Jie et al · 2018
Closest in time.
“Probabilistically distorted risk-sensitive infinite-horizon dynamic programming”
K. Lin, C. Jie and S.. Marcus · 2018
Closest in time.
“Stochastic variance-reduced policy gradient”
M. Papini et al · 2018
Closest in time.
“Adaptive system optimization using random directions stochastic approximation”
L.. Prashanth, S. Bhatnagar, M.. Fu and S.. Marcus · 2018
Closest in time.
“Reinforcement Learning: An Introduction”
Richard Sutton and Andrew Barto · 2018
Closest in time.
“Concentration of risk measures: A Wasserstein distance approach”
S.. Bhat and L.. Prashanth · 2019
Closest in time.
“Concentration bounds for empirical conditional value-at-risk: The unbounded case”
R.. Kolla, L.. Prashanth, S.. Bhat and K.. Jagannathan · 2019
Closest in time.
“Hessian aided policy gradient”
Z. Shen et al · 2019
Closest in time.
“Concentration Inequalities for Conditional Value at Risk”
P. Thomas and E. Learned-Miller · 2019
Closest in time.
“Concentration bounds for CVaR estimation: The cases of light-tailed and heavy-tailed distributions”
L.. Prashanth, K. Jagannathan and R.. Kolla · 2020
Closest in time.
“Global convergence of policy gradient methods to (almost) locally optimal policies”
K. Zhang, A. Koppel, H. Zhu and T. Basar · 2020
Closest in time.
“Stochastic optimization with momentum: Convergence, fluctuations, and traps avoidance”
A. Barakat, P. Bianchi, W. Hachem and S. Schechtman · 2021
Closest in time.
“A Policy Gradient Algorithm for the Risk-Sensitive Exponential Cost MDP”, 2022
M. Moharrami, Y. Murthy, A. Roy and R. Srikant · 2022
Closest in time.