Provably efficient imitation learning from observation alone
Sun, W., Vemula, A., Boots, B., and Bagnell, D. (2019) · 2019
Later among the works it cites.
Doubly robust bias reduction in infinite horizon off-policy estimation
Tang, Z., Feng, Y., Li, L., Zhou, D., and Liu, Q. (2019) · 2019
Later among the works it cites.
Is a good representation sufficient for sample efficient reinforcement learning?
Du, S. S., Kakade, S. M., Wang, R., and Yang, L. F. (2020) · 2020
Later among the works it cites.
A theoretical analysis of deep q-learning
Fan, J., Wang, Z., Xie, Y., and Yang, Z. (2020) · 2020
Later among the works it cites.
Near-optimal algorithms for minimax optimization
Lin, T., Jin, C., and Jordan, M. I. (2020) · 2020
Later among the works it cites.
Minimax Weight and Q-Function Learning for Off-Policy Evaluation
Uehara, M., Huang, J., and Jiang, N. (2020) · 2020
Later among the works it cites.
Randomized linear programming solves the markov decision problem in nearly linear (sometimes sublinear) time
Wang, M. (2020) · 2020
Later among the works it cites.
Q* approximation schemes for batch reinforcement learning: A theoretical comparison
Xie, T. and Jiang, N. (2020) · 2020
Later among the works it cites.
Infinite-horizon offline reinforcement learning with linear function approximation: Curse of dimensionality and algorithm
Original
Chen, L., Scherrer, B., and Bartlett, P. L. (2021) · 2021
Later among the works it cites.
Bilinear classes: A structural framework for provable generalization in rl
Du, S. S., Kakade, S. M., Lee, J. D., Lovett, S., Mahajan, G., Sun, W., and Wang, R. (2021) · 2021
Later among the works it cites.
Offline reinforcement learning: Fundamental barriers for value function approximation
Original
Foster, D. J., Krishnamurthy, A., Simchi-Levi, D., and Xu, Y. (2021) · 2021
Later among the works it cites.
Optidice: Offline policy optimization via stationary distribution correction estimation
Original
Lee, J., Jeon, W., Lee, B.-J., Pineau, J., and Kim, K.-E. (2021) · 2021
Later among the works it cites.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Rashidinejad, P., Zhu, B., Ma, C., Jiao, J., and Russell, S. (2021) · 2021
Later among the works it cites.
Pessimistic model-based offline rl: Pac bounds and posterior sampling under partial coverage
Original
Uehara, M. and Sun, W. (2021) · 2021
Later among the works it cites.
Bellman-consistent pessimism for offline reinforcement learning
Original
Xie, T., Cheng, C.-A., Jiang, N., Mineiro, P., and Agarwal, A. (2021) · 2021
Later among the works it cites.
Exponential lower bounds for batch reinforcement learning: Batch rl can be exponentially harder than online rl
Zanette, A. (2021) · 2021
Later among the works it cites.
Gendice: Generalized offline estimation of stationary values
Zhang, R., Dai, B., Li, L., and Schuurmans, D. (2020) · 2021
Later among the works it cites.