Fetching the paper…
Reading the bibliography…
This paper concerns the central issues of model robustness and sample efficiency in offline reinforcement learning (RL), which aims to learn to perform decision making from history data without active exploration.
Distributionally robust reinforcement learning
Smirnova, E., Dohmatob, E., and Mary, J. (2019) · 1902
Earlier work this paper cites.
Wasserstein robust reinforcement learning
Abdullah, M. A., Ren, H., Ammar, H. B., Milenkovic, V., Luo, R., Zhang, M., and Wang, J. (2019) · 1907
Earlier work this paper cites.
Distributionally robust optimization: A review
Rahimian, H. and Mehrotra, S. (2019) · 1908
Earlier work this paper cites.
A comparison of signalling alphabets
Gilbert, E. N. (1952) · 1952
Earlier work this paper cites.
Fast Bellman updates for robust MDPs
Ho, C. P., Petrik, M., and Wiesemann, W. (2018) · 1988
Earlier work this paper cites.
Error bounds for approximate value iteration
Munos, R. (2005) · 1999
Earlier work this paper cites.
Distributional robustness and regularization in reinforcement learning
Derman, E. and Mannor, S. (2020) · 2003
Earlier work this paper cites.
Robustness in markov decision problems with uncertain transition matrices
Nilim, A. and Ghaoui, L. (2003) · 2003
Earlier work this paper cites.
Robust dynamic programming
Iyengar, G. N. (2005) · 2005
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J. (2020) · 2005
Earlier work this paper cites.
Robust control of Markov decision processes with uncertain transition matrices
Nilim, A. and El Ghaoui, L. (2005) · 2005
Earlier work this paper cites.
Robust reinforcement learning with Wasserstein constraint
Hou, L., Pang, L., Hong, X., Lan, Y., Ma, Z., and Yin, D. (2020) · 2006
Earlier work this paper cites.
Optimistic distributionally robust policy optimization
Song, J. and Zhao, C. (2020) · 2006
Earlier work this paper cites.
Some bounds for the logarithmic function
Topsøe, F. (2007) · 2007
Earlier work this paper cites.
Reinforcement learning and savings behavior
Choi, J. J., Laibson, D., Madrian, B. C., and Metrick, A. (2009) · 2009
Earlier work this paper cites.
Introduction to nonparametric estimation
Tsybakov, A. B. and Zaiats, V. (2009) · 2009
Earlier work this paper cites.
Distributionally robust optimization under moment uncertainty with application to data-driven problems
Delage, E. and Ye, Y. (2010) · 2010
Earlier work this paper cites.
Robust control of uncertain markov decision processes with temporal logic specifications
Wolff, E. M., Topcu, U., and Murray, R. M. (2012) · 2012
Earlier work this paper cites.
Distributionally robust Markov decision processes
Xu, H. and Mannor, S. (2012) · 2012
Earlier work this paper cites.
Kullback-leibler divergence constrained distributionally robust optimization
Hu, Z. and Hong, L. J. (2013) · 2013
Earlier work this paper cites.
Robust modified policy iteration
Kaufman, D. L. and Schaefer, A. J. (2013) · 2013
Cited alongside, same era.
Finding locally optimal, collision-free trajectories with sequential convex optimization
Schulman, J., Ho, J., Lee, A. X., Awwal, I., Bradlow, H., and Abbeel, P. (2013) · 2013
Cited alongside, same era.
Robust markov decision processes
Wiesemann, W., Kuhn, D., and Rustem, B. (2013) · 2013
Cited alongside, same era.
Scaling up robust MDPs using function approximation
Tamar, A., Mannor, S., and Xu, H. (2014) · 2014
Cited alongside, same era.
A comprehensive survey on safe reinforcement learning
Garcıa, J. and Fernández, F. (2015) · 2015
Cited alongside, same era.
Reinforcement learning under model mismatch
Roy, A., Xu, H., and Pokutta, S. (2017) · 2017
Cited alongside, same era.
Partial policy iteration for l1-robust markov decision processes
Ho, C. P., Petrik, M., and Wiesemann, W. (2021) · 2021
Later among the works it cites.
Is pessimism provably efficient for offline RL?
Jin, Y., Yang, Z., and Wang, Z. (2021) · 2021
Later among the works it cites.
Sample complexity of asynchronous Q-learning: Sharper analysis and variance reduction
Li, G., Wei, Y., Chi, Y., Gu, Y., and Chen, Y. (2021) · 2021
Later among the works it cites.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Rashidinejad, P., Zhu, B., Ma, C., Jiao, J., and Russell, S. (2021) · 2021
Later among the works it cites.
Online robust reinforcement learning with model uncertainty
Wang, Y. and Zou, S. (2021) · 2021
Later among the works it cites.
Policy finetuning: Bridging sample-efficient offline and online reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A convex optimization approach to distributionally robust markov decision processes with Wasserstein distance
Yang, I. (2017) · 2017
Cited alongside, same era.
Data-driven robust optimization
Bertsimas, D., Gupta, V., and Kallus, N. (2018) · 2018
Cited alongside, same era.
Introductory lectures on stochastic optimization
Duchi, J. C. (2018) · 2018
Cited alongside, same era.
Certifying some distributional robustness with principled adversarial training
Sinha, A., Namkoong, H., and Duchi, J. (2018) · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Cited alongside, same era.
High-dimensional probability: An introduction with applications in data science
Vershynin, R. (2018) · 2018
Cited alongside, same era.
Xie, T., Jiang, N., Wang, H., Xiong, C., and Bai, Y. (2021) · 2021
Later among the works it cites.
Near-optimal offline reinforcement learning via double variance reduction
Yin, M., Bai, Y., and Wang, Y.-X. (2021) · 2021
Later among the works it cites.
Yin, M. and Wang, Y.-X. (2021) · 2021
Later among the works it cites.
Finite-sample regret bound for distributionally robust offline tabular reinforcement learning
Zhou, Z., Bai, Q., Zhou, Z., Qiu, L., Blanchet, J., and Glynn, P. (2021) · 2021
Later among the works it cites.
Finite-sample guarantees for Wasserstein distributionally robust optimization: Breaking the curse of dimensionality
Gao, R. (2022) · 2022
Closest in time.
Robust markov decision processes: Beyond rectangularity
Goyal, V. and Grand-Clement, J. (2022) · 2022
Closest in time.
Settling the sample complexity of model-based offline reinforcement learning
Li, G., Shi, L., Chen, Y., Chi, Y., and Wei, Y. (2022) · 2022
Closest in time.
Sample complexity of robust reinforcement learning with a generative model
Panaganti, K. and Kalathil, D. (2022) · 2022
Closest in time.
Pessimistic Q-learning for offline reinforcement learning: Towards optimal sample complexity
Shi, L., Li, G., Wei, Y., Chen, Y., and Chi, Y. (2022) · 2022
Closest in time.
Reliable off-policy evaluation for reinforcement learning
Wang, J., Gao, R., and Zha, H. (2022) · 2022
Closest in time.
Toward theoretical understandings of robust markov decision processes: Sample complexity and asymptotics
Yang, W., Zhang, L., and Zhang, Z. (2022) · 2022
Closest in time.
Seeing is not believing: Robust reinforcement learning against spurious correlation
Ding, W., Shi, L., Chi, Y., and Zhao, D. (2023) · 2023
Closest in time.
The curious price of distributional robustness in reinforcement learning with a generative model
Shi, L., Li, G., Wei, Y., Chen, Y., Geist, M., and Chi, Y. (2023) · 2023
Closest in time.
The efficacy of pessimism in asynchronous Q-learning
Yan, Y., Li, G., Chen, Y., and Fan, J. (2023) · 2023
Closest in time.