Off-policy evaluation via off-policy classification
A. Irpan, K. Rao, K. Bousmalis, C. Harris, J. Ibarz, and S. Levine · 2019
Later among the works it cites.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Original
N. Jaques, A. Ghandeharioun, J. H. Shen, C. Ferguson, A. Lapedriza, N. Jones, S. Gu, and R. Picard · 2019
Later among the works it cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
A. Kumar, J. Fu, M. Soh, G. Tucker, and S. Levine · 2019
Later among the works it cites.
Batch policy learning under constraints
H. Le, C. Voloshin, and Y. Yue · 2019
Later among the works it cites.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
O. Nachum, Y. Chow, B. Dai, and L. Li · 2019
Later among the works it cites.
Making efficient use of demonstrations to solve hard exploration problems
Original
T. L. Paine, C. Gulcehre, B. Shahriari, M. Denil, M. Hoffman, H. Soyer, R. Tanburn, S. Kapturowski, N. Rabinowitz, D. Williams, et al · 2019
Later among the works it cites.
Minimax weight and q-function learning for off-policy evaluation
Original
M. Uehara and N. Jiang · 2019
Later among the works it cites.
Empirical study of off-policy policy evaluation for reinforcement learning
Original
C. Voloshin, H. M. Le, N. Jiang, and Y. Yue · 2019
Later among the works it cites.
D4rl: Datasets for deep data-driven reinforcement learning
Original
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Closest in time.
Popcorn: Partially observed prediction constrained reinforcement learning
Original
J. Futoma, M. C. Hughes, and F. Doshi-Velez · 2020
Closest in time.
Rl unplugged: Benchmarks for offline reinforcement learning
Original
C. Gulcehre, Z. Wang, A. Novikov, T. L. Paine, S. G. Colmenarejo, K. Zolna, R. Agarwal, J. Merel, D. Mankowitz, C. Paduraru, G. Dulac-Arnold, J. Li, M. Norouzi, M. Hoffman, N. Ofir, T. George, N. Heess, and N. de Freitas · 2020
Closest in time.
Acme: A research framework for distributed reinforcement learning
Original
M. Hoffman, B. Shahriari, J. Aslanides, G. Barth-Maron, F. Behbahani, T. Norman, A. Abdolmaleki, A. Cassirer, F. Yang, K. Baumli, S. Henderson, A. Novikov, S. G. Colmenarejo, S. Cabi, C. Gulcehre, T. L. Paine, A. Cowie, Z. Wang, B. Piot, and N. de Freitas · 2020
Closest in time.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Original
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Closest in time.
Keep doing what worked: Behavior modelling priors for offline reinforcement learning
N. Siegel, J. T. Springenberg, F. Berkenkamp, A. Abdolmaleki, M. Neunert, T. Lampe, R. Hafner, N. Heess, and M. Riedmiller · 2020
Closest in time.
dm_control: Software and tasks for continuous control
Original
Y. Tassa, S. Tunyasuvunakool, A. Muldal, Y. Doron, S. Liu, S. Bohez, J. Merel, T. Erez, T. Lillicrap, and N. Heess · 2020
Closest in time.
Critic regularized regression
Original
Z. Wang, A. Novikov, K. Żołna, J. T. Springenberg, S. Reed, B. Shahriari, N. Siegel, J. Merel, C. Gulcehre, N. Heess, and N. de Freitas · 2020
Closest in time.