Fetching the paper…
Reading the bibliography…
Our goal is to compute a policy that guarantees improved return over a baseline policy even when the available MDP model is inaccurate.
On the complexity of solving Markov decision problems
M. Littman, T. Dean, and L. Kaelbling · 1995
Earlier work this paper cites.
Constrained Markov decision processes
E. Altman · 1999
Earlier work this paper cites.
Nonlinear programming
D. Bertsekas · 1999
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
S. Kakade and J. Langford · 2002
Earlier work this paper cites.
Markov decision processes with uncertain transition rates: Sensitivity and robust control
S. Kalyanasundaram, E. Chong, and N. Shroff · 2002
Earlier work this paper cites.
Envelope theorems for arbitrary choice sets
P. Milgrom and I. Segal · 2002
Earlier work this paper cites.
Action elimination and stopping conditions for reinforcement learning
E. Even-Dar, S. Mannor, and Y. Mansour · 2003
Earlier work this paper cites.
Inequalities for the
T. Weissman, E. Ordentlich, G. Seroussi, S. Verdu, and M. Weinberger · 2003
Earlier work this paper cites.
Convex optimization
S. Boyd and L. Vandenberghe · 2004
Cited alongside, same era.
Robust dynamic programming
G. Iyengar · 2005
Cited alongside, same era.
Robust control of Markov decision processes with uncertain transition matrices
A. Nilim and L. El Ghaoui · 2005
Cited alongside, same era.
Robust dynamic programming for discounted infinite-horizon Markov decision processes with uncertain stationary transition matrices
B. Li and S. Si · 2007
Cited alongside, same era.
Robust, risk-sensitive, and data-driven control of Markov decision processes
Y. Le Tallec · 2007
Cited alongside, same era.
Distributionally robust optimization under moment uncertainty with application to data-driven problems
E. Delage and Y. Ye · 2010
Cited alongside, same era.
Regret based robust solutions for uncertain Markov decision processes
A. Ahmed, P. Varakantham, Y. Adulyasak, and P. Jaillet · 2013
Later among the works it cites.
Robust modified policy iteration
D. Kaufman and A. Schaefer · 2013
Later among the works it cites.
Robust Markov decision processes
W. Wiesemann, D. Kuhn, and B. Rustem · 2013
Later among the works it cites.
Two-stage robust integer programming
G. Hanasusanto, D. Kuhn, and W. Wiesemann · 2014
Later among the works it cites.
RAAM : The benefits of robustness in approximating aggregated MDPs in reinforcement learning
M. Petrik · 2014
Later among the works it cites.
Markov decision processes: discrete stochastic dynamic programming
M. Puterman · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Conservative and greedy approaches to classification-based policy iteration
M. Ghavamzadeh and A. Lazaric · 2012
Cited alongside, same era.
High confidence off-policy evaluation
P. Thomas, G. Theocharous, and M. Ghavamzadeh · 2015
Closest in time.