Fetching the paper…
Reading the bibliography…
Robust Reinforcement Learning tries to make predictions more robust to changes in the dynamics or rewards of the system.
Distributionally robust reinforcement learning
Smirnova, E., Dohmatob, E., and Mary, J · 1902
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
Auer, P · 2002
Earlier work this paper cites.
Scalable first-order methods for robust mdps
Grand-Clément, J. and Kroer, C · 2005
Earlier work this paper cites.
Robust reinforcement learning
Morimoto, J. and Doya, K · 2005
Earlier work this paper cites.
First-order methods for wasserstein distributionally robust mdp
Grand-Clément, J. and Kroer, C · 2009
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Robust markov decision processes
Wiesemann, W., Kuhn, D., and Rustem, B · 2013
Earlier work this paper cites.
Approximate modified policy iteration and its application to the game of tetris
Scherrer, B., Ghavamzadeh, M., Gabillon, V., Lesner, B., and Geist, M · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Moritz, P., Jordan, M. I., and Abbeel, P · 2015
Earlier work this paper cites.
Statistics of robust optimization: A generalized empirical likelihood approach
Duchi, J., Glynn, P., and Namkoong, H · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Earlier work this paper cites.
Distributional reinforcement learning with quantile regression
Dabney, W., Rowland, M., Bellemare, M. G., and Munos, R · 2017
Earlier work this paper cites.
A convex optimization approach to distributionally robust markov decision processes with wasserstein distance
Yang, I · 2017
Earlier work this paper cites.
Implicit quantile networks for distributional reinforcement learning
Dabney, W., Ostrovski, G., Silver, D., and Munos, R · 2018
Earlier work this paper cites.
Learning models with uniform performance via distributionally robust optimization
Duchi, J. and Namkoong, H · 2018
Cited alongside, same era.
Robust empirical optimization is almost the same as mean–variance optimization
Gotoh, J.-y., Kim, M. J., and Lim, A. E · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Wasserstein robust reinforcement learning
Abdullah, M. A., Ren, H., Ammar, H. B., Milenkovic, V., Luo, R., Zhang, M., and Wang, J · 2019
Cited alongside, same era.
A theory of regularized markov decision processes
Geist, M., Scherrer, B., and Pietquin, O · 2019
Cited alongside, same era.
Twice regularized mdps and the equivalence between robustness and regularization
Derman, E., Geist, M., and Mannor, S · 2021
Later among the works it cites.
Maximum entropy rl (provably) solves some robust rl problems
Eysenbach, B. and Levine, S · 2021
Later among the works it cites.
Regularized policies are reward robust
Husain, H., Ciosek, K., and Tomioka, R · 2021
Later among the works it cites.
Robust risk-aware reinforcement learning
Jaimungal, S., Pesenti, S. M., Wang, Y. S., and Tatsat, H · 2021
Later among the works it cites.
Conservative offline distributional reinforcement learning
Ma, Y. J., Jayaraman, D., and Bastani, O · 2021
Later among the works it cites.
Tactical optimism and pessimism for deep reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Beyond confidence regions: Tight bayesian ambiguity sets for robust mdps
Petrik, M. and Russel, R. H · 2019
Cited alongside, same era.
Stable baselines3, 2019
Raffin, A., Hill, A., Ernestus, M., Gleave, A., Kanervisto, A., and Dormann, N · 2019
Cited alongside, same era.
Action robust reinforcement learning and applications in continuous control
Tessler, C., Efroni, Y., and Mannor, S · 2019
Cited alongside, same era.
Distributional robustness and regularization in reinforcement learning
Derman, E. and Mannor, S · 2020
Cited alongside, same era.
Controlling overestimation bias with truncated mixture of continuous distributional quantile critics
Kuznetsov, A., Shvechikov, P., Grishin, A., and Vetrov, D · 2020
Cited alongside, same era.
Improving robustness via risk averse distributional reinforcement learning
Singh, R., Zhang, Q., and Chen, Y · 2020
Cited alongside, same era.
Leverage the average: an analysis of kl regularization in reinforcement learning
Vieillard, N., Kozuno, T., Scherrer, B., Pietquin, O., Munos, R., and Geist, M · 2020
Cited alongside, same era.
Moskovitz, T., Parker-Holder, J., Pacchiano, A., Arbel, M., and Jordan, M · 2021
Later among the works it cites.
Gmac: A distributional perspective on actor-critic framework
Nam, D. W., Kim, Y., and Park, C. Y · 2021
Later among the works it cites.
Risk-averse offline reinforcement learning
Urpí, N. A., Curi, S., and Krause, A · 2021
Later among the works it cites.
Towards safe reinforcement learning via constraining conditional value-at-risk
Ying, C., Zhou, X., Su, H., Yan, D., and Zhu, J · 2021
Later among the works it cites.
Safe distributional reinforcement learning
Zhang, J. and Weng, P · 2021
Later among the works it cites.
Pessimistic bootstrapping for uncertainty-driven offline reinforcement learning
Bai, C., Wang, L., Yang, Z., Deng, Z., Garg, A., Liu, P., and Wang, Z · 2022
Closest in time.
Your policy regularizer is secretly an adversary
Brekelmans, R., Genewein, T., Grau-Moya, J., Delétang, G., Kunesch, M., Legg, S., and Ortega, P · 2022
Closest in time.
Efficient policy iteration for robust markov decision processes via regularization
Kumar, N., Levy, K., Wang, K., and Mannor, S · 2022
Closest in time.