Is q-learning provably efficient?
Jin, C., Allen-Zhu, Z., Bubeck, S., & Jordan, M. I. (2018) · 2018
Later among the works it cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Liu, Q., Li, L., Tang, Z., & Zhou, D. (2018) · 2018
Later among the works it cites.
Deep reinforcement learning for vision-based robotic grasping: A simulated comparative evaluation of off-policy methods
Quillen, D., Jang, E., Nachum, O., Finn, C., Ibarz, J., & Levine, S. (2018) · 2018
Later among the works it cites.
Behaviour policy estimation in off-policy policy evaluation: Calibration matters
Original
Raghu, A., Gottesman, O., Liu, Y., Komorowski, M., Faisal, A., Doshi-Velez, F., & Brunskill, E. (2018) · 2018
Later among the works it cites.
Near-optimal time and sample complexities for solving markov decision processes with a generative model
Sidford, A., Wang, M., Wu, X., Yang, L., & Ye, Y. (2018) · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S., & Barto, A. G. (2018) · 2018
Later among the works it cites.
Information-theoretic considerations in batch reinforcement learning
Chen, J., & Jiang, N. (2019) · 2019
Later among the works it cites.
Guidelines for reinforcement learning in healthcare
Gottesman, O., Johansson, F., Komorowski, M., Faisal, A., Sontag, D., Doshi-Velez, F., & Celi, L. A. (2019) · 2019
Later among the works it cites.
Batch policy learning under constraints
Le, H., Voloshin, C., & Yue, Y. (2019) · 2019
Later among the works it cites.
Off-policy policy gradient with state distribution correction
Liu, Y., Swaminathan, A., Agarwal, A., & Brunskill, E. (2019) · 2019
Later among the works it cites.
Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
Xie, T., Ma, Y., & Wang, Y.-X. (2019) · 2019
Later among the works it cites.
Model-based reinforcement learning with a generative model is minimax optimal
Agarwal, A., Kakade, S., & Yang, L. F. (2020) · 2020
Closest in time.
Robonet: Large-scale multi-robot learning
Dasari, S., Ebert, F., Tian, S., Nair, S., Bucher, B., Schmeckpeper, K., Singh, S., Levine, S., & Finn, C. (2020) · 2020
Closest in time.
Minimax-optimal off-policy evaluation with linear function approximation
Duan, Y., Jia, Z., & Wang, M. (2020) · 2020
Closest in time.
Provably efficient reinforcement learning with linear function approximation
Jin, C., Yang, Z., Wang, Z., & Jordan, M. I. (2020) · 2020
Closest in time.
Asymptotically efficient off-policy evaluation for tabular reinforcement learning
Yin, M., & Wang, Y.-X. (2020) · 2020
Closest in time.