Linear feature encoding for reinforcement learning
Song, Z., Parr, R., Liao, X., and Carin, L. (2016) · 2016
Later among the works it cites.
An emphatic approach to the problem of off-policy temporal-difference learning
Sutton, R. S., Mahmood, A. R., and White, M. (2016) · 2016
Later among the works it cites.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, O. P., and Zaremba, W. (2017) · 2017
Later among the works it cites.
Successor features for transfer in reinforcement learning
Barreto, A., Dabney, W., Munos, R., Hunt, J. J., Schaul, T., van Hasselt, H. P., and Silver, D. (2017) · 2017
Later among the works it cites.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R. (2017) · 2017
Later among the works it cites.
Learning to act by predicting the future
Dosovitskiy, A. and Koltun, V. (2017) · 2017
Later among the works it cites.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K. (2017) · 2017
Later among the works it cites.
Shallow updates for deep reinforcement learning
Levine, N., Zahavy, T., Mankowitz, D., Tamar, A., and Mannor, S. (2017) · 2017
Later among the works it cites.
A Laplacian framework for option discovery in reinforcement learning
Machado, M. C., Bellemare, M. G., and Bowling, M. (2017) · 2017
Later among the works it cites.
Boosted fitted q-iteration
Tosatto, S., Pirotta, M., D’Eramo, C., and Restelli, M. (2017) · 2017
Later among the works it cites.
Feature selection by singular value decomposition for reinforcement learning
Behzadian, B. and Petrik, M. (2018) · 2018
Later among the works it cites.
Feature-based aggregation and deep reinforcement learning: A survey and some new implementations
Bertsekas, D. P. (2018) · 2018
Later among the works it cites.
Dopamine: A research framework for deep reinforcement learning
Castro, P. S., Moitra, S., Gelada, C., Kumar, S., and Bellemare, M. G. (2018) · 2018
Later among the works it cites.
Combined reinforcement learning via abstract representations
François-Lavet, V., Bengio, Y., Precup, D., and Pineau, J. (2018) · 2018
Later among the works it cites.
Eigenoption discovery through the deep successor representation
Machado, M. C., Rosenbaum, C., Guo, X., Liu, M., Tesauro, G., and Campbell, M. (2018) · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
van den Oord, A., Li, Y., and Vinyals, O. (2018) · 2018
Later among the works it cites.
Deep reinforcement learning and the deadly triad
van Hasselt, H., Doron, Y., Strub, F., Hessel, M., Sonnerat, N., and Modayil, J. (2018) · 2018
Later among the works it cites.
Two-timescale networks for nonlinear value function approximation
Chung, W., Nath, S., Joseph, A. G., and White, M. (2019) · 2019
Closest in time.
The value function polytope in reinforcement learning
Dadashi, R., Taïga, A. A., Roux, N. L., Schuurmans, D., and Bellemare, M. G. (2019) · 2019
Closest in time.
DeepMDP: Learning continuous latent space modelsfor representation learning
Gelada, C., Kumar, S., Buckman, J., Nachum, O., and Bellemare, M. G. (2019) · 2019
Closest in time.
Near-optimal representation learning for hierarchical reinforcement learning
Nachum, O., Gu, S., Lee, H., and Levine, S. (2019) · 2019
Closest in time.
An Atari model zoo for analyzing, visualizing, and comparing deep reinforcement learning agents
Such, F. P., Madhavan, V., Liu, R., Wang, R., Castro, P. S., Li, Y., Schubert, L., Bellemare, M. G., Clune, J., and Lehman, J. (2019) · 2019
Closest in time.
The laplacian in rl: Learning representations with efficient approximations
Wu, Y., Tucker, G., and Nachum, O. (2019) · 2019
Closest in time.