RL unplugged: Benchmarks for offline reinforcement learning
Original
C. Gulcehre, Z. Wang, A. Novikov, T. Le Paine, S. G. Colmenarejo, K. Zolna, R. Agarwal, J. Merel, D. Mankowitz, C. Paduraru, G. Dulac-Arnold, J. Li, M. Norouzi, M. Hoffman, O. Nachum, G. Tucker, N. Heess, and N. de Freitas · 2020
Later among the works it cites.
Convergence proof for actor-critic methods applied to PPO and RUDDER
Original
M. Holzleitner, L. Gruber, J. A. Arjona-Medina, J. Brandstetter, and S. Hochreiter · 2020
Later among the works it cites.
Towards continual reinforcement learning: A review and perspectives
Original
K. Khetarpal, M. Riemer, I. Rish, and D. Precup · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
Original
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Original
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Later among the works it cites.
Offline reinforcement learning hands-on
Original
L. Monier, J. Kmec, A. Laterre, T. Pierrot, V. Courgeau, O. Sigaud, and K. Beguir · 2020
Later among the works it cites.
A unifying view of optimism in episodic reinforcement learning
G. Neu and C. Pike-Burke · 2020
Later among the works it cites.
Align-RUDDER: Learning from few demonstrations by reward redistribution
Original
V. P. Patil, M. Hofmarcher, M.-C. Dinu, M. Dorfer, P. M. Blies, J. Brandstetter, J. A. Arjona-Medina, and S. Hochreiter · 2020
Later among the works it cites.
Rl-cyclegan: Reinforcement learning aware simulation-to-real
K. Rao, C. Harris, A. Irpan, S. Levine, J. Ibarz, and M. Khansari · 2020
Later among the works it cites.
Mdp homomorphic networks: Group symmetries in reinforcement learning
E. van der Pol, D. Worrall, H. van Hoof, F. Oliehoek, and M. Welling · 2020
Later among the works it cites.
Critic regularized regression
Original
Z. Wang, A. Novikov, K. Zolna, J. T. Springenberg, S. Reed, B. Shahriari, N. Siegel, J. Merel, C. Gulcehre, N. Heess, and N. de Freitas · 2020
Later among the works it cites.
Bdd100k: A diverse driving dataset for heterogeneous multitask learning
Original
F. Yu, H. Chen, X. Wang, W. Xian, Y. Chen, F. Liu, V. Madhavan, and T. Darrell · 2020
Later among the works it cites.
D4rl: Datasets for deep data-driven reinforcement learning
Original
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2021
Closest in time.
A minimalist approach to offline reinforcement learning
S. Fujimoto and S.S. Gu · 2021
Closest in time.
Regularized behavior value estimation
Original
C. Gulcehre, S. Gómez Colmenarejo, Z. Wang, J. Sygnowski, T. Paine, K. Zolna, Y. Chen, M. Hoffman, R. Pascanu, and N. de Freitas · 2021
Closest in time.
Offline reinforcement learning with fisher divergence critic regularization
I. Kostrikov, R. Fergus, J. Tompson, and O. Nachum · 2021
Closest in time.
Collect & infer - a fresh look at data-efficient reinforcement learning
Original
M. A. Riedmiller, J. T. Springenberg, R. Hafner, and N. Heess · 2021
Closest in time.
Modern Hopfield Networks for Return Decomposition for Delayed Rewards
M. Widrich, M. Hofmarcher, V. P. Patil, A. Bitto-Nemling, and S. Hochreiter · 2021
Closest in time.
A theory of abstraction in reinforcement learning
Original
D. Abel · 2022
Closest in time.
XAI and Strategy Extraction via Reward Redistribution , pages 177–205
M.-C. Dinu, M. Hofmarcher, V. P. Patil, M. Dorfer, P. M. Blies, J. Brandstetter, J. A. Arjona-Medina, and S. Hochreiter · 2022
Closest in time.
Comparing model-free and model-based algorithms for offline reinforcement learning
Original
P. Swazinna, S. Udluft, D. Hein, and T. Runkler · 2022
Closest in time.