Reinforcement learning: Theory and algorithms
Agarwal, A., Jiang, N., Kakade, S. M., and Sun, W. (2021) · 2021
Later among the works it cites.
Minimax sample complexity for turn-based stochastic game
Cui, Q. and Yang, L. F. (2021) · 2021
Later among the works it cites.
Umbrella: Uncertainty-aware model-based offline reinforcement learning leveraging planning
Original
Diehl, C., Sievernich, T., Krüger, M., Hoffmann, F., and Bertran, T. (2021) · 2021
Later among the works it cites.
Optimal policy evaluation using kernel-based temporal difference methods
Original
Duan, Y., Wang, M., and Wainwright, M. J. (2021) · 2021
Later among the works it cites.
Nearly minimax optimal reinforcement learning for discounted MDPs
He, J., Zhou, D., and Gu, Q. (2021) · 2021
Later among the works it cites.
Is pessimism provably efficient for offline RL?
Jin, Y., Yang, Z., and Wang, Z. (2021) · 2021
Later among the works it cites.
Breaking the sample complexity barrier to regret-optimal model-free reinforcement learning
Li, G., Shi, L., Chen, Y., Gu, Y., and Chi, Y. (2021) · 2021
Later among the works it cites.
Sample complexity of offline reinforcement learning with deep ReLU networks
Original
Nguyen-Tang, T., Gupta, S., and Venkatesh, S. (2021) · 2021
Later among the works it cites.
Nearly horizon-free offline reinforcement learning
Ren, T., Li, J., Dai, B., Du, S. S., and Sanghavi, S. (2021) · 2021
Later among the works it cites.
Model selection for offline reinforcement learning: Practical considerations for healthcare settings
Tang, S. and Wiens, J. (2021) · 2021
Later among the works it cites.
Pessimistic model-based offline reinforcement learning under partial coverage
Original
Uehara, M. and Sun, W. (2021) · 2021
Later among the works it cites.
Batch value-function approximation with only realizability
Xie, T. and Jiang, N. (2021) · 2021
Later among the works it cites.
Policy finetuning: Bridging sample-efficient offline and online reinforcement learning
Original
Xie, T., Jiang, N., Wang, H., Xiong, C., and Bai, Y. (2021) · 2021
Later among the works it cites.
A unified off-policy evaluation approach for general value function
Original
Xu, T., Yang, Z., Wang, Z., and Liang, Y. (2021) · 2021
Later among the works it cites.
Towards instance-optimal offline reinforcement learning with pessimism
Yin, M. and Wang, Y.-X. (2021) · 2021
Later among the works it cites.
Provable benefits of actor-critic methods for offline reinforcement learning
Zanette, A., Wainwright, M. J., and Brunskill, E. (2021) · 2021
Later among the works it cites.
Finite-sample regret bound for distributionally robust offline tabular reinforcement learning
Zhou, Z., Zhou, Z., Bai, Q., Qiu, L., Blanchet, J., and Glynn, P. (2021) · 2021
Later among the works it cites.
When is offline two-player zero-sum Markov game solvable?
Original
Cui, Q. and Du, S. S. (2022) · 2022
Closest in time.
Sample complexity of asynchronous Q-learning: Sharper analysis and variance reduction
Li, G., Wei, Y., Chi, Y., Gu, Y., and Chen, Y. (2022) · 2022
Closest in time.
Sample complexity of robust reinforcement learning with a generative model
Panaganti, K. and Kalathil, D. (2022) · 2022
Closest in time.
A survey on offline reinforcement learning: Taxonomy, review, and open problems
Original
Prudencio, R. F., Maximo, M. R., and Colombini, E. L. (2022) · 2022
Closest in time.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Rashidinejad, P., Zhu, B., Ma, C., Jiao, J., and Russell, S. (2022) · 2022
Closest in time.
Distributionally robust model-based offline reinforcement learning with near-optimal sample complexity
Original
Shi, L. and Chi, Y. (2022) · 2022
Closest in time.
Representation learning for online and offline RL in low-rank MDPs
Uehara, M., Zhang, X., and Sun, W. (2022) · 2022
Closest in time.
Model-based reinforcement learning is minimax-optimal for offline zero-sum Markov games
Original
Yan, Y., Li, G., Chen, Y., and Fan, J. (2022) · 2022
Closest in time.
Near-optimal offline reinforcement learning with linear representation: Leveraging variance information with pessimism
Yin, M., Duan, Y., Wang, M., and Wang, Y.-X. (2022) · 2022
Closest in time.
Offline reinforcement learning with realizability and single-policy concentrability
Original
Zhan, W., Huang, B., Huang, A., Jiang, N., and Lee, J. D. (2022) · 2022
Closest in time.
Pessimistic minimax value iteration: Provably efficient equilibrium learning from offline datasets
Original
Zhong, H., Xiong, W., Tan, J., Wang, L., Zhang, T., Wang, Z., and Yang, Z. (2022) · 2022
Closest in time.
Minimax-optimal reward-agnostic exploration in reinforcement learning
Original
Li, G., Yan, Y., Chen, Y., and Fan, J. (2023) · 2023
Closest in time.
The efficacy of pessimism in asynchronous Q-learning
Yan, Y., Li, G., Chen, Y., and Fan, J. (2023) · 2023
Closest in time.
Settling the sample complexity of online reinforcement learning
Original
Zhang, Z., Chen, Y., Lee, J. D., and Du, S. S. (2023) · 2023
Closest in time.