Fetching the paper…
Reading the bibliography…
We address the problem of inverse reinforcement learning in Markov decision processes where the agent is risk-sensitive.
H. Robbins and S. Monro, “A stochastic approximation method,” The Annals of Mathematical Statistics , vol. 22, no. 3, pp. 400–407, 1951
1951
Earlier work this paper cites.
D. Kahneman and A. Tversky, “Prospect theory: An analysis of decision under risk,” Econometrica , vol. 47, no. 2, pp. 263–291, 1979
1979
Earlier work this paper cites.
P. W. Millar, “Asymptotic minimax theorems for the sample distribution function,” Zeitschrift für Wahrscheinlichkeitstheorie und Verwandte Gebiete , vol. 48, no. 3, pp. 233–252, 1979
1979
Earlier work this paper cites.
A. Tversky and D. Kahneman, “The framing of decisions and the psychology of choice,” Science , vol. 211, no. 4481, pp. 453–458, Jan. 1981
1981
Earlier work this paper cites.
H. Robbins and D. Siegmund, A Convergence Theorem for Non Negative Almost Supermartingales and Some Applications . Springer New York, 1985, pp. 111–135
1985
Earlier work this paper cites.
——, “Rational choice and the framing of decisions,” J. Business , vol. 59, no. 4, pp. pp. S251–S278, 1986
1986
Earlier work this paper cites.
C. F. Camerer, “An experimental test of several generalized utility theories,” J. Risk and Uncertainty , vol. 2, no. 1, pp. 61–104, 1989
1989
Earlier work this paper cites.
P. Massart, “The tight constant in the dvoretzky-kiefer-wolfowitz inequality,” The Annals of Probability , vol. 18, pp. 1269–1283, 1990
1990
Earlier work this paper cites.
A. Tversky and D. Kahneman, “Loss aversion in riskless choice: A reference-dependent model,” The Quarterly J. Economics , vol. 106, no. 4, pp. 1039–1061, 1991
1991
Earlier work this paper cites.
A. Tversky and D. Kahneman, “Advances in prospect theory: Cumulative representation of uncertainty,” J. Risk and Uncertainty , vol. 5, no. 4, pp. 297–323, Oct 1992
1992
Earlier work this paper cites.
M. Heger, “Consideration of risk in reinforcement learning,” in Proc. 11th Inter. Conf. Machine Learning , 1994, pp. 105–111
1994
Earlier work this paper cites.
J. N. Tsitsiklis, “Asynchronous stochastic approximation and q-learning,” Machine Learning , vol. 16, no. 3, pp. 185–202, 1994
1994
Earlier work this paper cites.
J. Penot, “On the interchange of subdifferentiation and epi-convergence,” J. Mathematical Analysis and Applications , vol. 196, no. 2, pp. 676–698, 1995
1995
Earlier work this paper cites.
S.-I. Amari, “Natural gradient works efficiently in learning,” Neural Computation , vol. 10, no. 2, pp. 251–276, 1998
1998
Cited alongside, same era.
R. Gonzalez and G. Wu, “On the shape of the probability weighting function,” Cognitive Psychology , vol. 38, no. 1, pp. 129–166, 1999
1999
Cited alongside, same era.
P. Artzner, F. Delbaen, J.-M. Eber, and D. Heath, “Coherent measures of risk,” Mathematical Finance , vol. 9, no. 3, pp. 203–228, 1999
1999
Cited alongside, same era.
S. S. Sastry, Nonlinear Systems . Springer, 1999
1999
Cited alongside, same era.
A. Y. Ng and S. Russell, “Algorithms for Inverse Reinforcement Learning,” in Proc. 17th Inter. Conf. Machine Learning , 2000, pp. 663–670
2000
Cited alongside, same era.
J. Heinonen, “Lectures on Lipschitz Analysis,” 14th Jyväskylä Summer School , 2004
2004
Later among the works it cites.
P. Geibel and F. Wysotzki, “Risk-sensitive reinforcement learning applied to control under constraints,” J. Artificial Intelligence Research , vol. 24, pp. 81–108, 2005
2005
Later among the works it cites.
J. Peters, S. Vijayakumar, and S. Schaal, “Natural actor-critic,” in Proc. 16th European Conf. Machine Learning , 2005, pp. 280–291
2005
Later among the works it cites.
B. Köszegi and M. Rabin, “A model of reference-dependent preferences,” The Quarterly J. Economics , vol. 121, no. 4, pp. 1133–1165, 2006
2006
Later among the works it cites.
N. D. Ratliff, J. A. Bagnell, and M. A. Zinkevich, “Maximum margin planning,” in Proc. 23rd Inter. Conf. Machine Learning , 2006, pp. 729–736
2006
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Simon, “Bounded rationality in social science: Today and tomorrow,” Mind & Society , vol. 1, no. 1, pp. 25–39, Mar. 2000
2000
Cited alongside, same era.
S. P. Coraluppi and S. I. Marcus, “Mixed risk-neutral/minimax control of discrete-time, finite-state markov decision processes,” IEEE Trans. Autom. Control , vol. 45, no. 3, pp. 528–532, 2000
2000
Cited alongside, same era.
V. S. Borkar and S. P. Meyn, “Risk-sensitive optimal control for markov decision processes with monotone cost,” Mathematics of Operations Research , vol. 27, no. 1, pp. 192–209, 2002
2002
Cited alongside, same era.
O. Mihatsch and R. Neuneier, “Risk-sensitive reinforcement learning,” Machine Learning , vol. 49, no. 2, pp. 267–290, 2002
2002
Cited alongside, same era.
H. Föllmer and A. Schied, “Convex measures of risk and trading constraints,” Finance and Stochastics , vol. 6, no. 4, pp. 429–447, 2002
2002
Cited alongside, same era.
H. J. Kushner and G. G. Yin, Stochastic Approximation and Recursive Algorithms and Applications . Springer, 2003
2003
Cited alongside, same era.
A. Y. Kruger, “On Fréchet Subdifferentials,” J. Mathematical Sciences , vol. 116, no. 3, 2003
2003
Cited alongside, same era.
G. Neu and C. Szepesvári, “Apprenticeship learning using inverse reinforcement learning and gradient methods,” in Proc. 23rd Conf. Uncertainty in Artificial Intelligence , 2007, pp. 295–302
2007
Later among the works it cites.
A. J. Nagengast, D. A. Braun, and D. M. Wolpert, “Risk-sensitive optimal feedback control accounts for sensorimotor behavior under uncertainty,” PLOS Computational Biology , vol. 6, no. 7, pp. 1–15, 2010
2010
Later among the works it cites.
Y. Shen, W. Stannat, and K. Obermayer, “Risk-Sensitive Markov Control Processes,” SIAM J. Control Optimization , vol. 51, no. 5, pp. 3652–3672, 2013
2013
Later among the works it cites.
Y. Shen, M. J. Tobia, and K. Obermayer, “Risk-sensitive reinforcement learning,” Neural Computation , vol. 26, pp. 1298–1328, 2014
2014
Later among the works it cites.
A. Latif, Banach Contraction Principle and Its Generalizations . Springer International Publishing, 2014, pp. 33–64
2014
Later among the works it cites.
P. L.A., C. Jie, M. Fu, S. Marcus, and C. Szepesvári, “Cumulative prospect theory meets reinforcement learning: Prediction and control,” in Proc. 33rd Intern. Conf. on Machine Learning , vol. 48, 2016
2016
Later among the works it cites.
A. Majumdar, S. Singh, A. Mandlekar, and M. Provone, “Risk-sensitive inverse reinforcement learning via coherent risk models,” in Robotics: Science and Systems , 2017
2017
Closest in time.
E. Mazumdar, L. J. Ratliff, T. Fiez, and S. S. Sastry, “Gradient-based inverse risk-sensitive reinforcement learning,” in Proc. 56th IEEE Conf. Decision and Control , 2017
2017
Closest in time.