Revisiting the Arcade Learning Environment: Evaluation protocols and open problems for general agents
Machado, M. C., Bellemare, M. G., Talvitie, E., Veness, J., Hausknecht, M., and Bowling, M · 2018
Later among the works it cites.
Organizing experience: A deeper look at replay mechanisms for sample-based planning in continuous state domains
Pan, Y., Zaheer, M., White, A., Patterson, A., and White, M · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
Deep reinforcement learning and the deadly triad
van Hasselt, H., Doron, Y., Strub, F., Hessel, M., Sonnerat, N., and Modayil, J · 2018
Later among the works it cites.
Reconciling λ \lambda -returns with experience replay
Daley, B. and Amato, C · 2019
Later among the works it cites.
Hyperbolic discounting and learning over multiple horizons
Fedus, W., Gelada, C., Bengio, Y., Bellemare, M. G., and Larochelle, H · 2019
Later among the works it cites.
Diagnosing bottlenecks in deep Q-learning algorithms
Fu, J., Kumar, A., Soh, M., and Levine, S · 2019
Later among the works it cites.
Recurrent experience replay in distributed reinforcement learning
Kapturowski, S., Ostrovski, G., Dabney, W., Quan, J., and Munos, R · 2019
Later among the works it cites.
Sample-efficient deep reinforcement learning via episodic backward update
Lee, S. Y., Sungik, C., and Chung, S.-Y · 2019
Later among the works it cites.
Remember and forget for experience replay
Novati, G. and Koumoutsakos, P · 2019
Later among the works it cites.
Importance resampling for off-policy prediction
Schlegel, M., Chung, W., Graves, D., Qian, J., and White, M · 2019
Later among the works it cites.
When to use parametric models in reinforcement learning?
van Hasselt, H. P., Hessel, M., and Aslanides, J · 2019
Later among the works it cites.
Experience replay optimization
Zha, D., Lai, K.-H., Zhou, K., and Hu, X · 2019
Later among the works it cites.
An optimistic perspective on offline reinforcement learning
Agarwal, R., Schuurmans, D., and Norouzi, M · 2020
Closest in time.
Combining Q-learning and search with amortized value estimates
Hamrick, J. B., Bapst, V., Sanchez-Gonzalez, A., Pfaff, T., Weber, T., Buesing, L., and Battaglia, P. W · 2020
Closest in time.
Ranking policy gradient
Lin, K. and Zhou, J · 2020
Closest in time.
Attentive experience replay
Sun, P., Zhou, W., and Li, H · 2020
Closest in time.