Fetching the paper…
Reading the bibliography…
Robust MDPs (RMDPs) can be used to compute policies with provable worst-case guarantees in reinforcement learning.
Neuro-dynamic programming
Bertsekas, D. P. and Tsitsiklis, J. N · 1996
Earlier work this paper cites.
Reinforcement learning
Sutton, R. S. and Barto, A · 1998
Earlier work this paper cites.
Solving Uncertain Markov Decision Processes
Bagnell, J. A., Ng, A. Y., and Schneider, J. G · 2001
Earlier work this paper cites.
Markov decision processes with uncertain transition rates: Sensitivity and robust control
Kalyanasundaram, S., Chong, E. K. P., and Shroff, N. B · 2002
Earlier work this paper cites.
Inequalities for the L1 deviation of the empirical distribution
Weissman, T., Ordentlich, E., Seroussi, G., Verdu, S., and Weinberger, M. J · 2003
Earlier work this paper cites.
Robust dynamic programming
Iyengar, G. N · 2005
Earlier work this paper cites.
Robust control of Markov decision processes with uncertain transition matrices
Nilim, A. and El Ghaoui, L · 2005
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
Puterman, M. L · 2005
Earlier work this paper cites.
The robustness-performance tradeoff in Markov decision processes
Xu, H. and Mannor, S · 2006
Earlier work this paper cites.
Robust, Risk-Sensitive, and Data-driven Control of Markov Decision Processes
Le Tallec, Y · 2007
Earlier work this paper cites.
Probably Approximately Correct (PAC) Exploration in Reinforcement Learning
Strehl, A. L · 2007
Earlier work this paper cites.
An analysis of model-based Interval Estimation for Markov Decision Processes
Strehl, A. and Littman, M · 2008
Earlier work this paper cites.
Robust Optimization
Ben-Tal, A., El Ghaoui, L., and Nemirovski, A · 2009
Earlier work this paper cites.
Parametric regret in uncertain Markov decision processes
Xu, H. and Mannor, S · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Auer, P., Jaksch, T., and Ortner, R · 2010
Earlier work this paper cites.
Near-optimal Regret Bounds for Reinforcement Learning
Jaksch, T., Ortner, R., and Auer, P · 2010
Cited alongside, same era.
Batch Reinforcement Learning
Lange, S., Gabel, T., and Riedmiller, M · 2012
Cited alongside, same era.
Lightning does not strike twice: Robust MDPs with coupled uncertainty
Mannor, S., Mebel, O., and Xu, H · 2012
Cited alongside, same era.
Machine Learning: A Probabilistic Perspective
Murphy, K · 2012
Cited alongside, same era.
Approximate dynamic programming by minimizing distributionally robust bounds
Petrik, M · 2012
Cited alongside, same era.
PAC optimal planning for invasive species management: Improved exploration for reinforcement learning from simulator-defined MDPs
Dietterich, T., Taleghan, M., and Crowley, M · 2013
Cited alongside, same era.
Doubly Robust Off-policy Value Evaluation for Reinforcement Learning
Jiang, N. and Li, L · 2015
Later among the works it cites.
Toward Minimax Off-policy Value Estimation
Li, L., Munos, R., and Szepesvári, C · 2015
Later among the works it cites.
PAC Optimal MDP Planning with Application to Invasive Species Management
Taleghan, M. A., Dietterich, T. G., Crowley, M., Hall, K., and Albers, H. J · 2015
Later among the works it cites.
High Confidence Off-Policy Evaluation
Thomas, P. S., Teocharous, G., and Ghavamzadeh, M · 2015
Later among the works it cites.
Real-time dynamic programming for Markov decision processes with imprecise probabilities
Delgado, K. V., De Barros, L. N., Dias, D. B., and Sanner, S · 2016
Later among the works it cites.
Robust MDPs with k-rectangular uncertainty
Mannor, S., Mebel, O., and Xu, H · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Robust Data-Driven Dynamic Programming
Hanasusanto, G. and Kuhn, D · 2013
Cited alongside, same era.
Reinforcement Learning in Robust Markov Decision Processes
Lim, S. H., Xu, H., and Mannor, S · 2013
Cited alongside, same era.
Robust Markov decision processes
Wiesemann, W., Kuhn, D., and Rustem, B · 2013
Cited alongside, same era.
Bayesian Data Analysis
Gelman, A., Carlin, J. B., Stern, H. S., and Rubin, D. B · 2014
Cited alongside, same era.
RAAM : The benefits of robustness in approximating aggregated MDPs in reinforcement learning
Petrik, M. and Subramanian, D · 2014
Cited alongside, same era.
Lectures on stochastic programming: Modeling and theory
Shapiro, A., Dentcheva, D., and Ruszczynski, A · 2014
Cited alongside, same era.
Safe and Efficient Off-Policy Reinforcement Learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M. G · 2016
Later among the works it cites.
Safe Policy Improvement by Minimizing Robust Baseline Regret
Petrik, M., Mohammad Ghavamzadeh, and Chow, Y · 2016
Later among the works it cites.
Data-efficient off-policy policy evaluation for reinforcement learning
Thomas, P. S. and Brunskill, E · 2016
Later among the works it cites.
Data-driven robust optimization
Bertsimas, D., Kallus, N., and Gupta, V · 2017
Later among the works it cites.
Robust Markov Decision Process: Beyond Rectangularity
Goyal, V. and Grand-Clement, J · 2018
Later among the works it cites.
Fast Bellman Updates for Robust MDPs
Ho, C. P., Petrik, M., and Wiesemann, W · 2018
Later among the works it cites.
Safe Policy Improvement with Baseline Bootstrapping, 2018
Laroche, R. and Trichelair, P · 2018
Later among the works it cites.
Policy-Conditioned Uncertainty Sets for Robust Markov Decision Processes
Tirinzoni, A., Milano, P., Chen, X., and Ziebart, B. D · 2018
Later among the works it cites.