Fetching the paper…
Reading the bibliography…
We study finite episodic Markov decision processes incorporating dynamic risk measures to capture risk sensitivity.
Risk-sensitive markov decision processes
Howard, R. A. and Matheson, J. E. (1972) · 1972
Earlier work this paper cites.
Coherent measures of risk
Artzner, P., Delbaen, F., Eber, J.-M., and Heath, D. (1999) · 1999
Earlier work this paper cites.
Risk sensitive asset allocation
Bielecki, T. R., Pliska, S. R., and Sherris, M. (2000) · 2000
Earlier work this paper cites.
Optimization of conditional value-at-risk
Rockafellar, R. T., Uryasev, S., et al. (2000) · 2000
Earlier work this paper cites.
Distortion risk measures: Coherence and stochastic dominance
Wirch, J. L. and Hardy, M. R. (2001) · 2001
Earlier work this paper cites.
Spectral measures of risk: A coherent representation of subjective risk aversion
Acerbi, C. (2002) · 2002
Earlier work this paper cites.
On the coherence of expected shortfall
Acerbi, C. and Tasche, D. (2002) · 2002
Earlier work this paper cites.
Inequalities for the l1 deviation of the empirical distribution
Weissman, T., Ordentlich, E., Seroussi, G., Verdu, S., and Weinberger, M. J. (2003) · 2003
Earlier work this paper cites.
Clinical data based optimal sti strategies for hiv: a reinforcement learning approach
Ernst, D., Stan, G.-B., Goncalves, J., and Wehenkel, L. (2006) · 2006
Earlier work this paper cites.
An old-new concept of convex risk measures: the optimized certainty equivalent
Ben-Tal, A. and Teboulle, M. (2007) · 2007
Earlier work this paper cites.
Risk-sensitive benchmarked asset management
Davis, M. and Lleo, S. (2008) · 2008
Earlier work this paper cites.
Entropic risk constraints for utility maximization
Rudloff, B., Sass, J., and Wunderlich, R. (2008) · 2008
Earlier work this paper cites.
Properties of distortion risk measures
Balbás, A., Garrido, J., and Mayoral, S. (2009) · 2009
Earlier work this paper cites.
Percentile optimization for markov decision processes with parameter uncertainty
Delage, E. and Mannor, S. (2010) · 2010
Earlier work this paper cites.
Risk-averse dynamic programming for markov decision processes
Ruszczyński, A. (2010) · 2010
Cited alongside, same era.
Entropic risk measures: Coherence vs. convexity, model ambiguity and robust large deviations
Föllmer, H. and Knispel, T. (2011) · 2011
Cited alongside, same era.
Iterated risk measures for risk-sensitive markov decision processes with discounted cost
Osogami, T. (2012) · 2012
Cited alongside, same era.
Convex risk measures: Basic facts, law-invariance and beyond, asymptotics for large portfolios
Föllmer, H. and Knispel, T. (2013) · 2013
Cited alongside, same era.
Risk-sensitive markov control processes
Shen, Y., Stannat, W., and Obermayer, K. (2013) · 2013
Cited alongside, same era.
Markov decision processes with iterated coherent risk measures
Model-based reinforcement learning with value-targeted regression
Jia, Z., Yang, L., Szepesvari, C., and Wang, M. (2020) · 2020
Later among the works it cites.
Reinforcement learning with dynamic convex risk measures
Coache, A. and Jaimungal, S. (2021) · 2021
Later among the works it cites.
Episodic reinforcement learning in finite mdps: Minimax lower bounds revisited
Domingues, O. D., Ménard, P., Kaufmann, E., and Valko, M. (2021) · 2021
Later among the works it cites.
Exponential bellman equation and improved regret bounds for risk-sensitive reinforcement learning
Fei, Y., Yang, Z., Chen, Y., and Wang, Z. (2021) · 2021
Later among the works it cites.
Off-policy risk assessment in contextual bandits
Huang, A., Leqi, L., Lipton, Z., and Azizzadenesheli, K. (2021) · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chu, S. and Zhang, Y. (2014) · 2014
Cited alongside, same era.
A note on a new class of recursive utilities in markov decision processes
Asienkiewicz, H. and Jaśkiewicz, A. (2017) · 2017
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Azar, M. G., Osband, I., and Munos, R. (2017) · 2017
Cited alongside, same era.
Regret bounds for markov decision processes with recursive optimized certainty equivalents
Xu, W., Gao, X., and He, X. (2023) · 2017
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Cited alongside, same era.
Explore first, exploit next: The true shape of regret in bandit problems
Garivier, A., Ménard, P., and Stoltz, G. (2019) · 2019
Cited alongside, same era.
Non-asymptotic gap-dependent regret bounds for tabular mdps
Simchowitz, M. and Jamieson, K. G. (2019) · 2019
Cited alongside, same era.
Bastani, O., Ma, J. Y., Shen, E., and Xu, W. (2022) · 2022
Later among the works it cites.
Markov decision processes with recursive risk measures
Bäuerle, N. and Glauner, A. (2022) · 2022
Later among the works it cites.
Conditionally elicitable dynamic risk measures for deep reinforcement learning
Coache, A., Jaimungal, S., and Cartea, Á. (2022) · 2022
Later among the works it cites.
Cascaded gaps: Towards logarithmic regret for risk-sensitive reinforcement learning
Fei, Y. and Xu, R. (2022) · 2022
Later among the works it cites.
Bridging distributional and risk-sensitive reinforcement learning with provable regret bounds
Liang, H. and Luo, Z.-Q. (2022) · 2022
Later among the works it cites.
A wasserstein distance approach for concentration of empirical risk estimates
Prashanth, L. and Bhat, S. P. (2022) · 2022
Later among the works it cites.
Provably efficient risk-sensitive reinforcement learning: Iterated cvar and worst path
Du, Y., Wang, S., and Huang, L. (2023) · 2023
Closest in time.
Risk-aware reinforcement learning with coherent risk measures and non-linear function approximation
Lam, T., Verma, A., Low, B. K. H., and Jaillet, P. (2023) · 2023
Closest in time.
Near-minimax-optimal risk-sensitive reinforcement learning with cvar
Wang, K., Kallus, N., and Sun, W. (2023) · 2023
Closest in time.