Fetching the paper…
Reading the bibliography…
We study risk-sensitive reinforcement learning in episodic Markov decision processes with unknown transition kernels, where the goal is to optimize the total reward under the risk measure of exponential utility.
Portfolio selection
Harry Markowitz · 1952
Earlier work this paper cites.
Risk-sensitive Markov decision processes
Ronald A. Howard and James E. Matheson · 1972
Earlier work this paper cites.
Risk-sensitive Optimal Control
Peter Whittle · 1990
Earlier work this paper cites.
Risk-sensitive control on an infinite time horizon
Wendell H Fleming and William M McEneaney · 1995
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Steven J. Bradtke and Andrew G. Barto · 1996
Earlier work this paper cites.
Risk sensitive control of Markov processes in countable state space
Daniel Hernández-Hernández and Steven I. Marcus · 1996
Earlier work this paper cites.
Risk-sensitive optimal control of hidden Markov models: Structural results
Emmanuel Fernández-Gaucherand and Steven I. Marcus · 1997
Earlier work this paper cites.
Risk sensitive Markov decision processes
Steven I. Marcus, Emmanual Fernández-Gaucherand, Daniel Hernández-Hernandez, Stefano Coraluppi, and Pedram Fard · 1997
Earlier work this paper cites.
Constrained Markov Decision Processes
Eitan Altman · 1999
Earlier work this paper cites.
Risk-sensitive and minimax control of discrete-time, finite-state Markov decision processes
Stefano P. Coraluppi and Steven I. Marcus · 1999
Earlier work this paper cites.
Risk-sensitive control of discrete-time Markov processes with infinite horizon
Giovanni B. Di Masi and Lukasz Stettner · 1999
Earlier work this paper cites.
The vanishing discount approach in Markov chains with risk-sensitive criteria
Rolando Cavazos-Cadena and Emmanuel Fernández-Gaucherand · 2000
Earlier work this paper cites.
Nash equilibria in risk-sensitive dynamic games
Margriet B. Klompstra · 2000
Earlier work this paper cites.
A sensitivity formula for risk-sensitive cost and the actor-critic algorithm
Vivek S. Borkar · 2001
Earlier work this paper cites.
Q-learning for risk-sensitive control
Vivek S. Borkar · 2002
Earlier work this paper cites.
Risk-sensitive optimal control for Markov decision processes with monotone cost
Vivek S. Borkar and Sean P. Meyn · 2002
Earlier work this paper cites.
Risk-sensitive reinforcement learning
Oliver Mihatsch and Ralph Neuneier · 2002
Earlier work this paper cites.
Risk-sensitive and risk-neutral multiarmed bandits
Eric V. Denardo, Haechurl Park, and Uriel G. Rothblum · 2007
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Zero-sum risk-sensitive stochastic differential games
Arnab Basu and Mrinal K. Ghosh · 2012
Earlier work this paper cites.
Neural prediction errors reveal a risk-sensitive reinforcement-learning process in the human brain
Yael Niv, Jeffrey A. Edlund, Peter Dayan, and John P. O’Doherty · 2012
Earlier work this paper cites.
Robustness and risk-sensitivity in Markov decision processes
Takayuki Osogami · 2012
Cited alongside, same era.
Risk-aversion in multi-armed bandits
Amir Sani, Alessandro Lazaric, and Rémi Munos · 2012
Cited alongside, same era.
The multi-armed bandit, with constraints
Eric V. Denardo, Eugene A Feinberg, and Uriel G Rothblum · 2013
Cited alongside, same era.
Robust risk-averse stochastic multi-armed bandits
Odalric-Ambrym Maillard · 2013
Cited alongside, same era.
Risk-sensitive Markov control processes
Yun Shen, Wilhelm Stannat, and Klaus Obermayer · 2013
Cited alongside, same era.
Sample complexity of risk-averse bandit-arm selection
Jia Yuan Yu and Evdokia Nikolova · 2013
Cited alongside, same era.
Decomposition of uncertainty in bayesian deep learning for efficient and risk-sensitive learning
Stefan Depeweg, José Miguel Hernández-Lobato, Finale Doshi-Velez, and Steffen Udluft · 2017
Later among the works it cites.
Safety-aware algorithms for adversarial contextual bandit
Wen Sun, Debadeepta Dey, and Ashish Kapoor · 2017
Later among the works it cites.
A general approach to multi-armed bandits under risk criteria
Asaf Cassel, Shie Mannor, and Assaf Zeevi · 2018
Later among the works it cites.
Risk-sensitive reinforcement learning: A constrained optimization viewpoint
Michael Fu et al · 2018
Later among the works it cites.
Is Q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I. Jordan · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Arnab Basu and Mrinal Kanti Ghosh · 2014
Cited alongside, same era.
More risk-sensitive Markov decision processes
Nicole Bäuerle and Ulrich Rieder · 2014
Cited alongside, same era.
Algorithms for cvar optimization in mdps
Yinlam Chow and Mohammad Ghavamzadeh · 2014
Cited alongside, same era.
Stationary Markov perfect equilibria in risk sensitive stochastic overlapping generations models
Anna Jaśkiewicz and Andrzej S. Nowak · 2014
Cited alongside, same era.
Generalization and exploration via randomized value functions
Ian Osband, Benjamin Van Roy, and Zheng Wen · 2014
Cited alongside, same era.
Risk-sensitive reinforcement learning
Yun Shen, Michael J. Tobia, Tobias Sommer, and Klaus Obermayer · 2014
Cited alongside, same era.
Tor Lattimore and Csaba Szepesvári · 2018
Later among the works it cites.
A block coordinate ascent algorithm for mean-variance optimization
Tengyang Xie, Bo Liu, Yangyang Xu, Mohammad Ghavamzadeh, Yinlam Chow, Daoming Lyu, and Daesub Yoon · 2018
Later among the works it cites.
The vanishing discount approach in a class of zero-sum finite games with risk-sensitive average criterion
Rolando Cavazos-Cadena and Daniel Hernández-Hernández · 2019
Later among the works it cites.
Lyapunov-based safe policy optimization for continuous control
Yinlam Chow, Ofir Nachum, Aleksandra Faust, Edgar Duenez-Guzman, and Mohammad Ghavamzadeh · 2019
Later among the works it cites.
Estimating risk and uncertainty in deep reinforcement learning
William R Clements, Benoît-Marie Robaglia, Bastien Van Delft, Reda Bahi Slaoui, and Sébastien Toth · 2019
Later among the works it cites.
Epistemic risk-sensitive reinforcement learning
Hannes Eriksson and Christos Dimitrakakis · 2019
Later among the works it cites.
Model and algorithm for time-consistent risk-aware Markov games
Wenjie Huang, Pham Viet Hai, and William B. Haskell · 2019
Later among the works it cites.
Provably efficient reinforcement learning with linear function approximation
Chi Jin, Zhuoran Yang, Zhaoran Wang, and Michael I. Jordan · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M. Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H. Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Later among the works it cites.
Nonzero-sum risk-sensitive finite-horizon continuous-time stochastic games
Qingda Wei · 2019
Later among the works it cites.
Provable self-play algorithms for competitive reinforcement learning
Yu Bai and Chi Jin · 2020
Closest in time.
Provably efficient safe exploration via primal-dual policy optimization
Dongsheng Ding, Xiaohan Wei, Zhuoran Yang, Zhaoran Wang, and Mihailo R Jovanović · 2020
Closest in time.
Exploration-exploitation in constrained mdps
Yonathan Efroni, Shie Mannor, and Matteo Pirotta · 2020
Closest in time.
Shuang Qiu, Xiaohan Wei, Zhuoran Yang, Jieping Ye, and Zhaoran Wang · 2020
Closest in time.
Constrained upper confidence reinforcement learning
Liyuan Zheng and Lillian J Ratliff · 2020
Closest in time.