Fetching the paper…
Reading the bibliography…
The Robust Markov Decision Process (RMDP) framework focuses on designing control policies that are robust against the parameter uncertainties due to the mismatches between the simulator model and real-world settings.
Raam: The benefits of robustness in approximating aggregated mdps in reinforcement learning
Petrik, M. and Subramanian, D. (2014) · 1987
Earlier work this paper cites.
An upper bound on the loss from approximate optimal-value functions
Singh, S. P. and Yee, R. C. (1994) · 1994
Earlier work this paper cites.
Robust and optimal control
Zhou, K., Doyle, J. C., Glover, K., et al. (1996) · 1996
Earlier work this paper cites.
Q-learning for risk-sensitive control
Borkar, V. S. (2002) · 2002
Earlier work this paper cites.
Robust dynamic programming
Iyengar, G. N. (2005) · 2005
Earlier work this paper cites.
Robust control of Markov decision processes with uncertain transition matrices
Nilim, A. and El Ghaoui, L. (2005) · 2005
Earlier work this paper cites.
Empirical bernstein bounds and sample-variance penalization
Maurer, A. and Pontil, M. (2009) · 2009
Earlier work this paper cites.
Distributionally robust Markov decision processes
Xu, H. and Mannor, S. (2010) · 2010
Earlier work this paper cites.
Minimax PAC bounds on the sample complexity of reinforcement learning with a generative model
Azar, M. G., Munos, R., and Kappen, H. J. (2013) · 2013
Earlier work this paper cites.
A course in robust control theory: a convex approach
Dullerud, G. E. and Paganini, F. (2013) · 2013
Earlier work this paper cites.
Reinforcement learning in robust Markov decision processes
Lim, S. H., Xu, H., and Mannor, S. (2013) · 2013
Earlier work this paper cites.
Robust Markov decision processes
Wiesemann, W., Kuhn, D., and Rustem, B. (2013) · 2013
Earlier work this paper cites.
Scaling up robust mdps using function approximation
Tamar, A., Mannor, S., and Xu, H. (2014) · 2014
Cited alongside, same era.
Probability in High Dimension
van Handel, R. (2014) · 2014
Cited alongside, same era.
Distributionally robust counterpart in Markov decision processes
Yu, P. and Xu, H. (2015) · 2015
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Cited alongside, same era.
Empirical dynamic programming
Haskell, W. B., Jain, R., and Kalathil, D. (2016) · 2016
Cited alongside, same era.
Robust mdps with k-rectangular uncertainty
Mannor, S., Mebel, O., and Xu, H. (2016) · 2016
Cited alongside, same era.
Kernel-based reinforcement learning in robust Markov decision processes
Lim, S. H. and Autef, A. (2019) · 2019
Later among the works it cites.
Beyond confidence regions: Tight bayesian ambiguity sets for robust mdps
Russel, R. H. and Petrik, M. (2019) · 2019
Later among the works it cites.
Model-based reinforcement learning with a generative model is minimax optimal
Agarwal, A., Kakade, S., and Yang, L. F. (2020) · 2020
Later among the works it cites.
A bayesian approach to robust reinforcement learning
Derman, E., Mankowitz, D., Mann, T., and Mannor, S. (2020) · 2020
Later among the works it cites.
Breaking the sample size barrier in model-based reinforcement learning with a generative model
Li, G., Wei, Y., Chi, Y., Gu, Y., and Chen, Y. (2020) · 2020
Later among the works it cites.
Robust reinforcement learning for continuous control with model misspecification
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Robust adversarial reinforcement learning
Pinto, L., Davidson, J., Sukthankar, R., and Gupta, A. (2017) · 2017
Cited alongside, same era.
Reinforcement learning under model mismatch
Roy, A., Xu, H., and Pokutta, S. (2017) · 2017
Cited alongside, same era.
Soft-robust actor-critic policy-gradient
Derman, E., Mankowitz, D. J., Mann, T. A., and Mannor, S. (2018) · 2018
Cited alongside, same era.
Near-optimal time and sample complexities for solving markov decision processes with a generative model
Sidford, A., Wang, M., Wu, X., Yang, L. F., and Ye, Y. (2018) · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Cited alongside, same era.
High-Dimensional Probability: An Introduction with Applications in Data Science
Vershynin, R. (2018) · 2018
Cited alongside, same era.
Mankowitz, D. J., Levine, N., Jeong, R., Abdolmaleki, A., Springenberg, J. T., Shi, Y., Kay, J., Hester, T., Mann, T., and Riedmiller, M. (2020) · 2020
Later among the works it cites.
SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python
Virtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., van der Walt, S. J., Brett, M., Wilson, J., Millman, K. J., Mayorov, N., Nelson, A. R. J., Jones, E., Kern, R., Larson, E., Carey, C. J., Polat, İ., Feng, Y., Moore, E. W., VanderPlas, J., Laxalde, D., Perktold, J., Cimrman, R., Henriksen, I., Quintero, E. A., Harris, C. R., Archibald, A. M., Ribeiro, A. H., Pedregosa, F., van Mulbregt, P., and SciPy 1.0 Contributors (2020) · 2020
Later among the works it cites.
Policy optimization for H 2 {}_{\mbox{2}} linear control with H ∞ {}_{\mbox{{$\infty$}}} robustness guarantee: Implicit regularization and global convergence
Zhang, K., Hu, B., and Basar, T. (2020b) · 2020
Later among the works it cites.
Empirical Q-Value Iteration
Kalathil, D., Borkar, V. S., and Jain, R. (2021) · 2021
Closest in time.
Robust reinforcement learning using least squares policy iteration with provable performance guarantees
Panaganti, K. and Kalathil, D. (2021) · 2021
Closest in time.
Yang, W., Zhang, L., and Zhang, Z. (2021) · 2021
Closest in time.
Finite-sample regret bound for distributionally robust offline tabular reinforcement learning
Zhou, Z., Bai, Q., Zhou, Z., Qiu, L., Blanchet, J., and Glynn, P. (2021) · 2021
Closest in time.