Fetching the paper…
Reading the bibliography…
In this paper, we study gap-dependent regret guarantees for risk-sensitive reinforcement learning based on the entropic risk measure.
Epistemic risk-sensitive reinforcement learning
Eriksson, H. & Dimitrakakis, C. (2019) · 1906
Earlier work this paper cites.
On the landscape of synchronization networks: A perspective from nonconvex optimization
Ling, S., Xu, R., & Bandeira, A. S. (2019) · 1907
Earlier work this paper cites.
Risk-sensitive control on an infinite time horizon
Fleming, W. H. & McEneaney, W. M. (1995) · 1915
Earlier work this paper cites.
A behavioral model of rational choice
Simon, H. A. (1955) · 1955
Earlier work this paper cites.
Hidden integrality of sdp relaxations for sub-gaussian mixture models
Fei, Y. & Chen, Y. (2018b) · 1965
Earlier work this paper cites.
Risk-sensitive Markov decision processes
Howard, R. A. & Matheson, J. E. (1972) · 1972
Earlier work this paper cites.
Optimal stochastic linear systems with exponential performance criteria and their relation to deterministic differential games
Jacobson, D. (1973) · 1973
Earlier work this paper cites.
Risk-sensitive Optimal Control
Whittle, P. (1990) · 1990
Earlier work this paper cites.
Risk sensitive control of Markov processes in countable state space
Hernández-Hernández, D. & Marcus, S. I. (1996) · 1996
Earlier work this paper cites.
Risk-sensitive and minimax control of discrete-time, finite-state Markov decision processes
Coraluppi, S. P. & Marcus, S. I. (1999) · 1999
Earlier work this paper cites.
Risk-sensitive control of discrete-time Markov processes with infinite horizon
Di Masi, G. B. & Stettner, L. (1999) · 1999
Earlier work this paper cites.
The vanishing discount approach in Markov chains with risk-sensitive criteria
Cavazos-Cadena, R. & Fernández-Gaucherand, E. (2000) · 2000
Earlier work this paper cites.
A sensitivity formula for risk-sensitive cost and the actor-critic algorithm
Borkar, V. S. (2001) · 2001
Earlier work this paper cites.
Q-learning for risk-sensitive control
Borkar, V. S. (2002) · 2002
Earlier work this paper cites.
Risk-sensitive optimal control for Markov decision processes with monotone cost
Borkar, V. S. & Meyn, S. P. (2002) · 2002
Cited alongside, same era.
Risk-sensitive reinforcement learning
Mihatsch, O. & Neuneier, R. (2002) · 2002
Cited alongside, same era.
Prediction, learning, and games
Cesa-Bianchi, N. & Lugosi, G. (2006) · 2006
Cited alongside, same era.
Entropic risk measures: Coherence vs. convexity, model ambiguity and robust large deviations
Föllmer, H. & Knispel, T. (2011) · 2011
Cited alongside, same era.
Robustness
Hansen, L. P. & Sargent, T. J. (2011) · 2011
Cited alongside, same era.
Neural prediction errors reveal a risk-sensitive reinforcement-learning process in the human brain
Niv, Y., Edlund, J. A., Dayan, P., & O’Doherty, J. P. (2012) · 2012
Cited alongside, same era.
Information theoretic mpc for model-based reinforcement learning
Williams, G., Wagener, N., Goldfain, B., Drews, P., Rehg, J. M., Boots, B., & Theodorou, E. A. (2017) · 2017
Later among the works it cites.
Is Q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., & Jordan, M. I. (2018) · 2018
Later among the works it cites.
Entropic risk measure in policy search
Nass, D., Belousov, B., & Peters, J. (2019) · 2019
Later among the works it cites.
Non-asymptotic gap-dependent regret bounds for tabular mdps
Simchowitz, M. & Jamieson, K. G. (2019) · 2019
Later among the works it cites.
More data can expand the generalization gap between adversarially robust and standard models
Chen, L., Min, Y., Zhang, M., & Karbasi, A. (2020) · 2020
Later among the works it cites.
Achieving the bayes error rate in synchronization and block models by sdp, robustly
Fei, Y. & Chen, Y. (2020) · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Robustness and risk-sensitivity in Markov decision processes
Osogami, T. (2012) · 2012
Cited alongside, same era.
Thermodynamics as a theory of decision-making with information-processing costs
Ortega, P. A. & Braun, D. A. (2013) · 2013
Cited alongside, same era.
Risk-sensitive Markov control processes
Shen, Y., Stannat, W., & Obermayer, K. (2013) · 2013
Cited alongside, same era.
More risk-sensitive Markov decision processes
Bäuerle, N. & Rieder, U. (2014) · 2014
Cited alongside, same era.
Risk-sensitive reinforcement learning
Shen, Y., Tobia, M. J., Sommer, T., & Obermayer, K. (2014) · 2014
Cited alongside, same era.
Human decision-making under limited time
Ortega, P. A. & Stocker, A. A. (2016) · 2016
Cited alongside, same era.
Later among the works it cites.
Risk-sensitive reinforcement learning: Near-optimal risk-sample tradeoff in regret
Fei, Y., Yang, Z., Chen, Y., Wang, Z., & Xie, Q. (2020) · 2020
Later among the works it cites.
Bandit algorithms
Lattimore, T. & Szepesvári, C. (2020) · 2020
Later among the works it cites.
Deep neural tangent kernel and laplace kernel have the same rkhs
Chen, L. & Xu, S. (2021) · 2021
Later among the works it cites.
Logarithmic regret for reinforcement learning with linear function approximation
He, J., Zhou, D., & Gu, Q. (2021) · 2021
Later among the works it cites.
Convergence and alignment of gradient descent with random back propagation weights
Song, G., Xu, R., & Lafferty, J. (2021) · 2021
Later among the works it cites.
Meta learning in the continuous time limit
Xu, R., Chen, L., & Karbasi, A. (2021) · 2021
Later among the works it cites.
Q-learning with logarithmic regret
Yang, K., Yang, L., & Du, S. (2021) · 2021
Later among the works it cites.