Fetching the paper…
Reading the bibliography…
The literature on Inverse Reinforcement Learning (IRL) typically assumes that humans take actions in order to minimize the expected value of a cost function, i.e., that humans are risk neutral.
Princeton University Press
von Neumann J and Morgenstern O (1944) Theory of Games and Economic Behavior · 1944
Earlier work this paper cites.
Econometrica 21(4): 503–546
Allais M (1953) Le comportement de l’homme rationnel devant le risque: critique des postulats et axiomes de l’école américaine · 1953
Earlier work this paper cites.
The Quarterly Journal of Economics 75(4): 643–669
Ellsberg D (1961) Risk, ambiguity, and the savage axioms · 1961
Earlier work this paper cites.
Management Science 8(7): 356–369
Howard R and Matheson J (1972) Risk-sensitive Markov decision processes · 1972
Earlier work this paper cites.
Econometrica : 263–291
Kahneman D and Tversky A (1979) Prospect theory: An analysis of decision under risk · 1979
Earlier work this paper cites.
Journal of Economic Behavior & Organization 3(4): 323–343
Quiggin J (1982) A theory of anticipated utility · 1982
Earlier work this paper cites.
Econometrica 55(1): 95–115
Yaari ME (1987) The dual theory of choice under risk · 1987
Earlier work this paper cites.
Mathematics of Operations Research 14(1): 147–161
Filar JA, Kallenberg LCM and Lee HM (1989) Variance-penalized Markov decision processes · 1989
Earlier work this paper cites.
Journal of Mathematical Economics 18(2): 141–153
Gilboa I and Schmeidler D (1989) Maxmin expected utility with non-unique prior · 1989
Earlier work this paper cites.
In: Proc. Computational Learning Theory
Russell S (1998) Learning agents for uncertain environments · 1998
Earlier work this paper cites.
Mathematical Finance 9(3): 203–228
Artzner P, Delbaen F, Eber JM and Heath D (1999) Coherent measures of risk · 1999
Earlier work this paper cites.
Journal of Mathematical Analysis and Applications 231(1): 47–67
Wu C and Yuanlie L (1999) Minimizing risk models in markov decision process with policies depending on target values · 1999
Earlier work this paper cites.
In: Int. Conf. on Machine Learning
Ng A and Russell S (2000) Algorithms for inverse reinforcement learning · 2000
Earlier work this paper cites.
Econometrica 68(5): 1281–1292
Rabin M (2000) Risk aversion and expected-utility theory: A calibration theorem · 2000
Earlier work this paper cites.
Journal of Risk 2: 21–41
Rockafellar RT and Uryasev S (2000) Optimization of conditional value-at-risk · 2000
Earlier work this paper cites.
Journal of Banking & Finance 26(7): 1505–1518
Acerbi C (2002) Spectral measures of risk: A coherent representation of subjective risk aversion · 2002
Earlier work this paper cites.
Journal of Banking & Finance 26(7): 1487–1503
Acerbi C and Tasche D (2002) On the coherence of expected shortfall · 2002
Earlier work this paper cites.
Machine Learning 49(2): 267–290
Mihatsch O and Neuneier R (2002) Risk-sensitive reinforcement learning · 2002
Earlier work this paper cites.
Journal of Banking & Finance 26(7): 1443–1471
Rockafellar RT and Uryasev S (2002) Conditional value-at-risk for general loss distributions · 2002
Earlier work this paper cites.
In: Int. Conf. on Machine Learning
Abbeel P and Ng AY (2004) Apprenticeship learning via inverse reinforcement learning · 2004
Earlier work this paper cites.
In: IEEE Int. Symp. on Computer Aided Control Systems Design
Löfberg J (2004) YALMIP : A toolbox for modeling and optimization in MATLAB · 2004
Earlier work this paper cites.
In: Int. Conf. on Machine Learning
Abbeel P and Ng AY (2005) Exploration and apprenticeship learning in reinforcement learning · 2005
Earlier work this paper cites.
SIAM Journal on Optimization 15(3): 751–779
Burke JV, Lewis AS and Overton ML (2005) A robust gradient sampling algorithm for nonsmooth nonconvex optimization · 2005
Earlier work this paper cites.
SIAM Journal on Optimization 16(1): 69–95
Eichhorn A and Römisch W (2005) Polyhedral risk measures in stochastic programming · 2005
Cited alongside, same era.
Journal of Artificial Intelligence Research 24(1): 81–108
Geibel P and Wysotzki F (2005) Risk-sensitive reinforcement learning applied to control under constraints · 2005
Cited alongside, same era.
Science 310(5754): 1680–1683
Hsu M, Bhatt M, Adolphs R, Tranel D and Camerer CF (2005) Neural systems responding to degrees of uncertainty in human decision-making · 2005
Cited alongside, same era.
Operations Research 53(5): 780–798
Nilim A and El Ghaoui L (2005) Robust control of Markov decision processes with uncertain transition matrices · 2005
Cited alongside, same era.
In: Advances in Neural Information Processing Systems
Kolter JZ, Abbeel P and Ng AY (2007) Hierarchical apprenticeship learning with application to quadruped locomotion · 2007
Cited alongside, same era.
In: International Joint Conference on Artificial Intelligence
In: Robotics: Science and Systems Workshop on Inverse Optimal Control and Robotic Learning from Demonstration
Park T and Levine S (2013) Inverse optimal control for humanoid locomotion · 2013
Later among the works it cites.
In: American Control Conference
Chow Y and Pavone M (2014) A framework for time-consistent, risk-averse model predictive control: Theory and algorithms · 2014
Later among the works it cites.
Second edition. Elsevier
Glimcher P and Fehr E (2014) Neuroeconomics · 2014
Later among the works it cites.
Econometrica 82(1): 1–39
Gul F and Pesendorfer W (2014) Expected uncertain utility theory · 2014
Later among the works it cites.
Second edition. SIAM
Shapiro A, Dentcheva D and Ruszczyński A (2014) Lectures on stochastic programming: Modeling and theory · 2014
Later among the works it cites.
Neural Computation 26(7): 1298–1328
Shen Y, Tobia MJ, Sommer T and Obermayer K (2014) Risk-sensitive reinforcement learning · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ramachandran D and Amir E (2007) Bayesian inverse reinforcement learning · 2007
Cited alongside, same era.
In: OR Tools and Applications: Glimpses of Future Technologies , chapter 3. INFORMS
Rockafellar RT (2007) Coherent approaches to risk in optimization under uncertainty · 2007
Cited alongside, same era.
In: Proc. AAAI Conf. on Artificial Intelligence
Ziebart BD, Maas A, Bagnell JA and Dey AK (2008) Maximum entropy inverse reinforcement learning · 2008
Cited alongside, same era.
Operations Research Letters 37(3): 143–147
Shapiro A (2009) On a time consistency concept in risk averse multi-stage stochastic programming · 2009
Cited alongside, same era.
In: IEEE/RSJ Int. Conf. on Intelligent Robots & Systems
Ziebart BD, Ratliff N, Gallagher G, Mertz C, Peterson K, Bagnell JA, Hebert M, Key AK and Srinivasa S (2009) Planning-based prediction for pedestrians · 2009
Cited alongside, same era.
Journal of Risk and Uncertainty 41(2): 81–111
Hey JD, Lotito G and Maffioletti A (2010) The descriptive and predictive adequacy of theories of decision making under uncertainty/ambiguity · 2010
Cited alongside, same era.
Autonomous Robots 28(3): 369–383
Mombaur K, Truong A and Laumond JP (2010) From human to humanoid locomotion—an inverse optimal control approach · 2010
Cited alongside, same era.
In: Advances in Neural Information Processing Systems
Chow Y, Tamar A, Mannor S and Pavone M (2015) Risk-sensitive and robust decision-making: a CVaR optimization approach · 2015
Later among the works it cites.
In: Int. Symp. on Robotics Research
Englert P and Toussaint M (2015) Inverse KKT – learning cost functions of manipulation tasks from demonstrations · 2015
Later among the works it cites.
In: Proc. IEEE Conf. on Robotics and Automation
Kuderer M, Gulati S and Burgard W (2015) Learning driving styles for autonomous vehicles from demonstration · 2015
Later among the works it cites.
Available at: https://arxiv.org/abs/1507.04888
Wulfmeier M, Ondruska P and Posner I (2015) Maximum entropy deep inverse reinforcement learning · 2015
Later among the works it cites.
In: IEEE Conf. on Decision and Control
Axelrod A, Carlone L, Chowdhary G and Karaman S (2016) Data-driven prediction of EVAR with confidence in time-varying datasets · 2016
Later among the works it cites.
PLoS ONE 11(12): e0167021
Carton D, Nitsch V, Meinzer D and Wollherr D (2016) Towards assessing the human trajectory planning horizon · 2016
Later among the works it cites.
In: Int. Conf. on Machine Learning
Finn C, Levine S and Abbeel P (2016) Guided cost learning: Deep inverse optimal control via policy optimization · 2016
Later among the works it cites.
In: Readings in Formal Epistemology , first edition, chapter 21
Gilboa I and Marinacci M (2016) Ambiguity and the Bayesian paradigm · 2016
Later among the works it cites.
Int. Journal of Robotics Research 35(11): 1289–1307
Kretzschmar H, Spies M, Sprunk C and Burgard W (2016) Socially compliant mobile robot navigation via inverse reinforcement learning · 2016
Later among the works it cites.
In: Int. Conf. on Machine Learning
Prashanth LA, Jie C, Fu M, Marcus S and Szepesvári C (2016) Cumulative prospect theory meets reinforcement learning: Prediction and control · 2016
Later among the works it cites.
IEEE Transactions on Automatic Control 62(7): 3323–3338
Tamar A, Chow Y, Ghavamzadeh M and Mannor S (2016) Sequential decision making with coherent risk · 2016
Later among the works it cites.
Available at https://mosek.com/
ApS M (2017) MOSEK optimization software · 2017
Closest in time.
Numerische Mathematik 136(2): 343–381
Lanza A, Morigi S, Selesnick I and Sgallari F (2017) Nonconvex nonsmooth optimization via convex–nonconvex majorization–minimization · 2017
Closest in time.
In: Int. Symp. on Robotics Research
Majumdar A and Pavone M (2017) How should a robot assess risk? Towards an axiomatic theory of risk in robotics · 2017
Closest in time.
In: Robotics: Science and Systems
Majumdar A, Singh S, Mandlekar A and Pavone M (2017) Risk-sensitive inverse reinforcement learning via coherent risk models · 2017
Closest in time.
Available at: https://arxiv.org/abs/1703.09842
Ratliff LJ and Mazumdar E (2017) Risk-sensitive inverse reinforcement learning via gradient methods · 2017
Closest in time.
Available at https://vires.com/vtd-vires-virtual-test-drive/
VIRES Simulationstechnologie GmbH (2017) VTD - Virtual Test Drive · 2017
Closest in time.