Fetching the paper…
Reading the bibliography…
In reinforcement learning, the standard criterion to evaluate policies in a state is the expectation of (discounted) sum of rewards.
An analog of the minimax theorem for vector payoffs
D. Blackwell · 1956
Earlier work this paper cites.
An axiomatic characterization of skew-symmetric bilinear functionals, with applications to utility theory
P.C. Fishburn · 1981
Earlier work this paper cites.
Percentiles and Markovian decision processes
Jerzy A. Filar · 1983
Earlier work this paper cites.
Utility, probabilistic constraints, mean and variance of discounted rewards in Markov decision processes
D. J. White · 1987
Earlier work this paper cites.
Variance-penalized Markov decision processes
Jerzy A. Filar, L. C. M. Kallenberg, and Huey-Miin Lee · 1989
Earlier work this paper cites.
Rationality and dynamic choice: Foundational explorations
E. McClennen · 1990
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
M.L. Puterman · 1994
Earlier work this paper cites.
Temporal difference learning and td-gammon
Gerald Tesauro · 1995
Earlier work this paper cites.
Stochastic approximation with time scales
V.S. Borkar · 1997
Earlier work this paper cites.
Implementing resolute choice under uncertainty
Jean-Yves Jaffray · 1998
Earlier work this paper cites.
Optimization models for the first arrival target distribution function in discrete time
Stella X. Yu, Yuanlie Lin, and Pingfan Yan · 1998
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
A.Y. Ng and S. Russell · 2000
Earlier work this paper cites.
Spoken dialogue management as planning and acting under uncertainty
B. Zhang, Q. Cai, J. Mao, E Chang, and B. Guo · 2001
Earlier work this paper cites.
Quantile Regression
R. Koenker · 2005
Earlier work this paper cites.
Risk-sensitive planning with one-switch utility functions: Value iteration
Y. Liu and S. Koenig · 2005
Earlier work this paper cites.
Value-at-Risk: The New Benchmark for Managing Financial Risk
Philippe Jorion · 2006
Earlier work this paper cites.
Functional value iteration for decision-theoretic planning with general utility functions
Y. Liu and S. Koenig · 2006
Cited alongside, same era.
Dynamic programming analysis of the TV game who wants to be a millionaire?
F. Perea and J. Puerto · 2006
Cited alongside, same era.
Dynamo: amazon’s highly available key-value store
G. DeCandia, D. Hastorun, M. Jampani, G. Kakulapati, A. Lakshman, A. Pilchin, S. Sivasubramanian, P. Vosshall, and W. Vogels · 2007
Cited alongside, same era.
Percentile optimization in uncertain Markov decision processes with application to efficient exploration
E. Delage and S. Mannor · 2007
Cited alongside, same era.
Pure stationary optimal strategies in Markov decision processes
Hugo Gimbert · 2007
Cited alongside, same era.
Stochastic approximation : a dynamical systems viewpoint
Vivek S. Borkar · 2008
Ordinal decision models for Markov decision processes
P. Weng · 2012
Later among the works it cites.
Preference-based reinforcement learning
R. Busa-Fekete, B. Szörenyi, P. Weng, W. Cheng, and E. Hüllermeier · 2013
Later among the works it cites.
Interactive value iteration for Markov decision processes with unknown rewards
P. Weng and B. Zanuttini · 2013
Later among the works it cites.
Sample complexity of risk-averse bandit-arm selection
Jia Yuan Yu and Evdokia Nikolova · 2013
Later among the works it cites.
Risk-constrained Markov decision processes
V. Borkar and Rahul Jain · 2014
Later among the works it cites.
Meet your expectations with guarantees: beyond worst-case synthesis in quantitative games
Véronique Bruyère, Emmanuel Filiot, Mickael Randour, and Jean-François Raskin · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Benefits of quantile regression for the analysis of customer lifetime value in a contractual setting: An application in financial services
D.F. Benoit and D. Van den Poel · 2009
Cited alongside, same era.
Regret based reward elicitation for Markov decision processes
K. Regan and C. Boutilier · 2009
Cited alongside, same era.
Autonomous helicopter aerobatics through apprenticeship learning
Pieter Abbeel, Adam Coates, and Andrew Y. Ng · 2010
Cited alongside, same era.
Markov Decision Processes in Artificial Intelligence , chapter Non-Standard Criteria, pages 319–359
M. Boussard, M. Bouzid, A.I. Mouaddib, R. Sabbadin, and P. Weng · 2010
Cited alongside, same era.
Quantile maximization in decision theory
M.J. Rostek · 2010
Cited alongside, same era.
Markov decision processes with average value-at-risk criteria
Nicole Bäuerle and Jonathan Ott · 2011
Cited alongside, same era.
Preference-based Reinforcement Learning: Evolutionary Direct Policy Search using a Preference-based Racing Algorithm
Robert Busa-Fekete, Balazs Szorenyi, Paul Weng, Weiwei Cheng, and Eyke Hüllermeier · 2014
Later among the works it cites.
Extreme bandits
Alexandra Carpentier and Michal Valko · 2014
Later among the works it cites.
Algorithms for cvar optimization in MDPs
Yinlam Chow and Mohammad Ghavamzadeh · 2014
Later among the works it cites.
A learning scheme for blackwell’s approachability in mdps and stackelberg stochastic games
D. Kalathil, V.S. Borkar, and R. Jain · 2014
Later among the works it cites.
Percentile queries in multi-dimensional Markov decision processes
Mickael Randour, Jean-François Raskin, and Ocan Sankur · 2014
Later among the works it cites.
QPRED: Using quantile predictions to improve power usage for private clouds
R. Wolski and J. Brevik · 2014
Later among the works it cites.
Solving MDPs with skew symmetric bilinear utility functions
Hugo Gilbert, Olivier Spanjaard, Paolo Viappiani, and Paul Weng · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Later among the works it cites.
Qualitative multi-armed bandits: A quantile-based approach
Balázs Szörényi, Róbert Busa-Fekete, Paul Weng, and Eyke Hüllermeier · 2015
Later among the works it cites.
Model-free reinforcement learning with skew-symmetric bilinear utilities
Hugo Gilbert, Bruno Zanuttini, Paolo Viappiani, Paul Weng, and Esther Nicart · 2016
Closest in time.