Fetching the paper…
Reading the bibliography…
An important problem in sequential decision-making under uncertainty is to use limited data to compute a safe policy, i.e., a policy that is guaranteed to perform at least as well as a given baseline strategy.
Approximately optimal approximate reinforcement learning
S. Kakade and J. Langford · 2002
Earlier work this paper cites.
Inequalities for the L 1 L_{1} deviation of the empirical distribution
T. Weissman, E. Ordentlich, G. Seroussi, S. Verdu, and M. Weinberger · 2003
Earlier work this paper cites.
Robust dynamic programming
G. Iyengar · 2005
Earlier work this paper cites.
Robust control of Markov decision processes with uncertain transition matrices
A. Nilim and L. El Ghaoui · 2005
Earlier work this paper cites.
Robust, Risk-Sensitive, and Data-driven Control of Markov Decision Processes
Y. Le Tallec · 2007
Earlier work this paper cites.
Conservative and greedy approaches to classification-based policy iteration
M. Ghavamzadeh and A. Lazaric · 2012
Cited alongside, same era.
Regret based Robust Solutions for Uncertain Markov Decision Processes
A. Ahmed and P Varakantham · 2013
Cited alongside, same era.
Strategy iteration is strongly polynomial for 2-player turn-based stochastic games with a constant discount factor
T. Hansen, P. Miltersen, and U. Zwick · 2013
Cited alongside, same era.
Safe Policy Iteration
M. Pirotta, M. Restelli, and D. Calandriello · 2013
Cited alongside, same era.
Robust Markov decision processes
W. Wiesemann, D. Kuhn, and B. Rustem · 2013
Cited alongside, same era.
RAAM : The benefits of robustness in approximating aggregated MDPs in reinforcement learning
M. Petrik and D. Subramanian · 2014
Later among the works it cites.
Off-policy model-based learning under unknown factored dynamics
A. Hallak, F. Schnitzler, T. Mann, and S. Mannor · 2015
Later among the works it cites.
Optimal Threshold Control for Energy Arbitrage with Degradable Battery Storage
M. Petrik and X. Wu · 2015
Later among the works it cites.
High confidence off-policy evaluation
P. Thomas, G. Theocharous, and M. Ghavamzadeh · 2015
Later among the works it cites.
High confidence policy improvement
P. Thomas, G. Theocharous, and M. Ghavamzadeh · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…