Fetching the paper…
Reading the bibliography…
To mitigate the limitation that the classical reinforcement learning (RL) framework heavily relies on identical training and test environments, Distributionally Robust RL (DRRL) has been proposed to enhance performance across a range of environments, possibly including unknown test environments.
Wasserstein Robust Reinforcement Learning, 2019
Abdullah, M. A., Ren, H., Ammar, H. B., Milenkovic, V., Luo, R., Zhang, M., and Wang, J · 1907
Earlier work this paper cites.
Option pricing: A simplified approach
Cox, J. C., Ross, S. A., and Rubinstein, M · 1979
Earlier work this paper cites.
Multinomial goodness-of-fit tests
Cressie, N. and Read, T. R · 1984
Earlier work this paper cites.
Asynchronous stochastic approximation and q-learning
Tsitsiklis, J. N · 1994
Earlier work this paper cites.
The ode method for convergence of stochastic approximation and reinforcement learning
Borkar, V. S. and Meyn, S. P · 2000
Earlier work this paper cites.
Bias and variance in value function estimation
Mannor, S., Simester, D., Sun, P., and Tsitsiklis, J. N · 2004
Earlier work this paper cites.
Robust dynamic programming
Iyengar, G. N · 2005
Earlier work this paper cites.
Robust control of markov decision processes with uncertain transition matrices
Nilim, A. and El Ghaoui, L · 2005
Earlier work this paper cites.
Stochastic approximation: a dynamical systems viewpoint , volume 48
Borkar, V. S · 2009
Earlier work this paper cites.
Distributionally robust markov decision processes
Xu, H. and Mannor, S · 2010
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Reinforcement learning in robust markov decision processes
Lim, S. H., Xu, H., and Mannor, S · 2013
Earlier work this paper cites.
Robust markov decision processes
Wiesemann, W., Kuhn, D., and Rustem, B · 2013
Earlier work this paper cites.
Scaling up robust mdps using function approximation
Tamar, A., Mannor, S., and Xu, H · 2014
Cited alongside, same era.
Unbiased monte carlo for optimization and functions of expectations via multi-level randomization
Blanchet, J. H. and Glynn, P. W · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al · 2016
Cited alongside, same era.
Model-free risk-sensitive reinforcement learning
Delétang, G., Grau-Moya, J., Kunesch, M., Genewein, T., Brekelmans, R., Legg, S., and Ortega, P. A · 2021
Later among the works it cites.
Learning models with uniform performance via distributionally robust optimization
Duchi, J. C. and Namkoong, H · 2021
Later among the works it cites.
Partial policy iteration for l1-robust markov decision processes
Ho, C. P., Petrik, M., and Wiesemann, W · 2021
Later among the works it cites.
Online robust reinforcement learning with model uncertainty
Wang, Y. and Zou, S · 2021
Later among the works it cites.
Finite-sample regret bound for distributionally robust offline tabular reinforcement learning
Zhou, Z., Zhou, Z., Bai, Q., Qiu, L., Blanchet, J., and Glynn, P · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforcement learning under model mismatch
Roy, A., Xu, H., and Pokutta, S · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Cited alongside, same era.
Maximum a posteriori policy optimisation
Abdolmaleki, A., Springenberg, J. T., Tassa, Y., Munos, R., Heess, N., and Riedmiller, M · 2018
Cited alongside, same era.
Soft-robust actor-critic policy-gradient
Derman, E., Mankowitz, D. J., Mann, T. A., and Mannor, S · 2018
Cited alongside, same era.
Wasserstein distributionally robust stochastic control: A data-driven approach
Yang, I · 2018
Cited alongside, same era.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., Oh, J., Horgan, D., Kroiss, M., Danihelka, I., Huang, A., Sifre, L., Cai, T., Agapiou, J. P., Jaderberg, M., Vezhnevets, A. S., Leblond, R., Pohlen, T., Dalibard, V., Budden, D., Sulsky, Y., Molloy, J., Paine, T. L., Gulcehre, C., Wang, Z., Pfaff, T., Wu, Y., Ring, R., Yogatama, D., Wünsch, D., McKinney, K., Smith, O., Schaul, T., Lillicrap, T., Kavukcuoglu, K., Hassabis, D., Apps, C., and Silver, D · 2019
Cited alongside, same era.
Robust reinforcement learning using least squares policy iteration with provable performance guarantees
Badrinath, K. P. and Kalathil, D · 2021
Cited alongside, same era.
Goyal, V. and Grand-Clement, J · 2022
Later among the works it cites.
Distributionally robust q q -learning
Liu, Z., Bai, Q., Blanchet, J., Dong, P., Xu, W., Zhou, Z., and Zhou, Z · 2022
Later among the works it cites.
Distributionally robust offline reinforcement learning with linear function approximation
Ma, X., Liang, Z., Xia, L., Zhang, J., Blanchet, J., Liu, M., Zhao, Q., and Zhou, Z · 2022
Later among the works it cites.
Robust q-learning algorithm for markov decision processes under wasserstein uncertainty
Neufeld, A. and Sester, J · 2022
Later among the works it cites.
Sample complexity of robust reinforcement learning with a generative model
Panaganti, K. and Kalathil, D · 2022
Later among the works it cites.
Robust reinforcement learning using offline data
Panaganti, K., Xu, Z., Kalathil, D., and Ghavamzadeh, M · 2022
Later among the works it cites.
Toward theoretical understandings of robust markov decision processes: Sample complexity and asymptotics
Yang, W., Zhang, L., and Zhang, Z · 2022
Later among the works it cites.