Fetching the paper…
Reading the bibliography…
We develop a stochastic approximation-type algorithm to solve finite state/action, infinite-horizon, risk-aware Markov decision processes.
On a stochastic approximation method
Kai Lai Chung · 1954
Earlier work this paper cites.
A markovian decision process
Richard Bellman · 1957
Earlier work this paper cites.
Weighted sums of certain dependent random variables
Kazuoki Azuma · 1967
Earlier work this paper cites.
Convex Analysis
R. Tyrrell Rockafellar · 1970
Earlier work this paper cites.
Probability, volume 7 of classics in applied mathematics
Leo Breiman · 1992
Earlier work this paper cites.
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
On the convergence of subdifferentials of convex functions
Hédy Attouch and Gerald Beer · 1993
Earlier work this paper cites.
On the convergence of subdifferentials of convex functions
Jean-Paul Penot · 1993
Earlier work this paper cites.
On the convergence of stochastic iterative dynamic programming algorithms
Tommi Jaakkola, Michael I Jordan, and Satinder P Singh · 1994
Earlier work this paper cites.
Asynchronous stochastic approximation and q-learning
John N Tsitsiklis · 1994
Earlier work this paper cites.
Neuro-dynamic programming
Dimitri P Bertsekas and John N Tsitsiklis · 1996
Earlier work this paper cites.
Variational Analysis
R. T. Rockafellar and R.J-B. Wets · 1998
Earlier work this paper cites.
Coherent measures of risk
Philippe Artzner, Freddy Delbaen, Jean-Marc Eber, and David Heath · 1999
Earlier work this paper cites.
The ode method for convergence of stochastic approximation and reinforcement learning
Vivek S Borkar and Sean P Meyn · 2000
Earlier work this paper cites.
Stability of locally optimal solutions
Adam B Levy, Ren6 A Poliquin, and R Tyrrell Rockafellar · 2000
Earlier work this paper cites.
Optimization of conditional value-at-risk
R Tyrrell Rockafellar and Stanislav Uryasev · 2000
Earlier work this paper cites.
Spectral measures of risk: a coherent representation of subjective risk aversion
Carlo Acerbi · 2002
Earlier work this paper cites.
Q-learning for risk-sensitive control
Vivek S Borkar · 2002
Earlier work this paper cites.
Convex measures of risk and trading constraints
Hans Föllmer and Alexander Schied · 2002
Earlier work this paper cites.
Portfolio optimization with conditional value-at-risk objective and constraints
Pavlo Krokhmal, Jonas Palmquist, and Stanislav Uryasev · 2002
Earlier work this paper cites.
Minimax analysis of stochastic problems
Alexander Shapiro and Anton Kleywegt · 2002
Earlier work this paper cites.
Nash q-learning for general-sum stochastic games
Junling Hu and Michael P Wellman · 2003
Earlier work this paper cites.
Stochastic Approximation and Recursive Algorithms and Applications
Harold J. Kushner and G.George Yin · 2003
Earlier work this paper cites.
Convex Optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Learning rates for q-learning
Eyal Even-Dar and Yishay Mansour · 2004
Cited alongside, same era.
On a class of minimax stochastic programs
Alexander Shapiro and Shabbir Ahmed · 2004
Cited alongside, same era.
An efficient stochastic approximation algorithm for stochastic saddle point problems
Arkadi Nemirovski and Reuven Rubinstein · 2005
Cited alongside, same era.
On duality theory of convex semi-infinite programming
Alexander Shapiro · 2005
Cited alongside, same era.
Have we met? mdp based speaker id for robot dialogue
Filip Krsmanovic, Curtis Spencer, Daniel Jurafsky, and Andrew Y Ng · 2006
Cited alongside, same era.
Optimization of convex risk functions
Andrzej Ruszczynski and Alexander Shapiro · 2006
Cited alongside, same era.
Minimax and risk averse multistage stochastic programming
Alexander Shapiro · 2012
Later among the works it cites.
More risk-sensitive markov decision processes
Nicole Bäuerle and Ulrich Rieder · 2013
Later among the works it cites.
Robust solutions of optimization problems affected by uncertain probabilities
Aharon Ben-Tal, Dick Den Hertog, Anja De Waegenaere, Bertrand Melenberg, and Gijs Rennen · 2013
Later among the works it cites.
Stochastic dominance-constrained markov decision processes
William B. Haskell and Rahul Jain · 2013
Later among the works it cites.
Training and evaluation of an mdp model for social multi-user human-robot interaction
Simon Keizer, Mary Ellen Foster, Oliver Lemon, Andre Gaschler, and Manuel Giuliani · 2013
Later among the works it cites.
Optimization with multivariate conditional value-at-risk constraints
Nilay Noyan and Gábor Rudolf · 2013
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
An old-new concept of convex risk measures: The optimized certainty equivalent
Aharon Ben-Tal and Marc Teboulle · 2007
Cited alongside, same era.
Markov decision processes with their applications
Qiying Hu and Wuyi Yue · 2007
Cited alongside, same era.
Approximate Dynamic Programming: Solving the curses of dimensionality
Warren B Powell · 2007
Cited alongside, same era.
Stochastic approximation
Vivek S Borkar et al · 2008
Cited alongside, same era.
Computing var and cvar using stochastic approximation and adaptive unconstrained importance sampling
Olivier Bardou, Noufel Frikha, and Gilles Pages · 2009
Cited alongside, same era.
Constructing uncertainty sets for robust linear optimization
Dimitris Bertsimas and David B Brown · 2009
Cited alongside, same era.
Later among the works it cites.
On solving multistage stochastic programs with coherent risk measures
Andy Philpott, Vitor de Matos, and Erlon Finardi · 2013
Later among the works it cites.
On kusuoka representation of law invariant risk measures
Alexander Shapiro · 2013
Later among the works it cites.
Risk-sensitive markov control processes
Yun Shen, Wilhelm Stannat, and Klaus Obermayer · 2013
Later among the works it cites.
Risk-constrained markov decision processes
Vivek Borkar and Rahul Jain · 2014
Later among the works it cites.
Policy gradients for cvar-constrained mdps
LA Prashanth · 2014
Later among the works it cites.
Risk-sensitive reinforcement learning
Yun Shen, Michael J Tobia, Tobias Sommer, and Klaus Obermayer · 2014
Later among the works it cites.
Policy gradients beyond expectations: Conditional value-at-risk
Aviv Tamar, Yonatan Glassner, and Shie Mannor · 2014
Later among the works it cites.
Markov decision processes with applications in wireless sensor networks: A survey
Mohammad Abu Alsheikh, Dinh Thai Hoang, Dusit Niyato, Hwee-Pink Tan, and Shaowei Lin · 2015
Later among the works it cites.
Deriving robust counterparts of nonlinear uncertain inequalities
Aharon Ben-Tal, Dick Den Hertog, and Jean-Philippe Vial · 2015
Later among the works it cites.
A convex analytic approach to risk-aware markov decision processes
William B Haskell and Rahul Jain · 2015
Later among the works it cites.
Kusuoka representations of coherent risk measures in general probability spaces
Nilay Noyan and Gábor Rudolf · 2015
Later among the works it cites.
Continuity of optimal solution functions and their conditions on objective functions
Yasushi Terazono and Ayumu Matani · 2015
Later among the works it cites.
Addressing environment non-stationarity by repeating q-learning updates
Sherief Abdallah and Michael Kaisers · 2016
Later among the works it cites.
Computationally tractable counterparts of distributionally robust constraints on risk measures
Krzysztof Postek, Dick den Hertog, and Bertrand Melenberg · 2016
Later among the works it cites.
Deep reinforcement learning with double q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Later among the works it cites.
Risk-averse approximate dynamic programming with quantile-based risk measures
Daniel R Jiang and Warren B Powell · 2017
Later among the works it cites.
Risk-constrained reinforcement learning with percentile risk criteria
Yinlam Chow, Mohammad Ghavamzadeh, Lucas Janson, and Marco Pavone · 2018
Closest in time.