Provably good batch reinforcement learning without great exploration
Liu, Y., Swaminathan, A., Agarwal, A., and Brunskill, E · 2020
Later among the works it cites.
Hyperparameter selection for offline reinforcement learning
Original
Paine, T. L., Paduraru, C., Michi, A., Gulcehre, C., Zolna, K., Novikov, A., Wang, Z., and de Freitas, N · 2020
Later among the works it cites.
A game theoretic framework for model based reinforcement learning
Rajeswaran, A., Mordatch, I., and Kumar, V · 2020
Later among the works it cites.
Q* approximation schemes for batch reinforcement learning: A theoretical comparison
Xie, T. and Jiang, N · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J. Y., Levine, S., Finn, C., and Ma, T · 2020
Later among the works it cites.
On the theory of policy gradient methods: Optimality, approximation, and distribution shift
Agarwal, A., Kakade, S. M., Lee, J. D., and Mahajan, G · 2021
Later among the works it cites.
Heuristic-guided reinforcement learning
Cheng, C.-A., Kolobov, A., and Swaminathan, A · 2021
Later among the works it cites.
A minimalist approach to offline reinforcement learning
Fujimoto, S. and Gu, S. S · 2021
Later among the works it cites.
Is pessimism provably efficient for offline rl?
Jin, Y., Yang, Z., and Wang, Z · 2021
Later among the works it cites.
Offline reinforcement learning with implicit q-learning
Original
Kostrikov, I., Nair, A., and Levine, S · 2021
Later among the works it cites.
Of moments and matching: A game-theoretic framework for closing the imitation gap
Swamy, G., Choudhury, S., Bagnell, J. A., and Wu, S · 2021
Later among the works it cites.
Representation learning for online and offline rl in low-rank mdps
Original
Uehara, M., Zhang, X., and Sun, W · 2021
Later among the works it cites.
A convergent and efficient deep q network algorithm
Original
Wang, Z. T. and Ueda, M · 2021
Later among the works it cites.
Bellman-consistent pessimism for offline reinforcement learning
Xie, T., Cheng, C.-A., Jiang, N., Mineiro, P., and Agarwal, A · 2021
Later among the works it cites.
Combo: Conservative offline model-based policy optimization
Yu, T., Kumar, A., Rafailov, R., Rajeswaran, A., Levine, S., and Finn, C · 2021
Later among the works it cites.
Provable benefits of actor-critic methods for offline reinforcement learning
Zanette, A., Wainwright, M. J., and Brunskill, E · 2021
Later among the works it cites.
Towards hyperparameter-free policy selection for offline reinforcement learning
Zhang, S. and Jiang, N · 2021
Later among the works it cites.
Stackelberg actor-critic: Game-theoretic reinforcement learning algorithms
Original
Zheng, L., Fiez, T., Alumbaugh, Z., Chasnov, B., and Ratliff, L. J · 2021
Later among the works it cites.