Fetching the paper…
Reading the bibliography…
We introduce a value-based RL agent, which we call BBF, that achieves super-human performance in the Atari 100K benchmark.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Python reference manual
Van Rossum, G. and Drake Jr, F. L · 1995
Earlier work this paper cites.
Introduction to Reinforcement Learning
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Bias-variance error bounds for temporal difference updates
Kearns, M. J. and Singh, S · 2000
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2005
Earlier work this paper cites.
Matplotlib: A 2d graphics environment
Hunter, J. D · 2007
Earlier work this paper cites.
Python for scientific computing
Oliphant, T. E · 2007
Earlier work this paper cites.
Mastering atari games with limited data
Ye, W., Liu, S., Kurutach, T., Abbeel, P., and Gao, Y · 2007
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Earlier work this paper cites.
How to discount deep reinforcement learning: Towards new dynamic strategies
François-Lavet, V., Fonteneau, R., and Ernst, D · 2015
Earlier work this paper cites.
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
van Hasselt, H., Guez, A., and Silver, D · 2016
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., Hasselt, H., Lanctot, M., and Freitas, N · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Earlier work this paper cites.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D · 2017
Earlier work this paper cites.
Jax: composable transformations of python+ numpy programs
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., et al · 2018
Earlier work this paper cites.
Dopamine: A Research Framework for Deep Reinforcement Learning
Castro, P. S., Moitra, S., Gelada, C., Kumar, S., and Bellemare, M. G · 2018
Earlier work this paper cites.
Implicit quantile networks for distributional reinforcement learning
Dabney, W., Ostrovski, G., Silver, D., and Munos, R · 2018
Earlier work this paper cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., et al · 2018
Earlier work this paper cites.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., Hasselt, H. V., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M. G., and Silver, D · 2018
Cited alongside, same era.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Machado, M. C., Bellemare, M. G., Talvitie, E., Veness, J., Hausknecht, M., and Bowling, M · 2018
Cited alongside, same era.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., et al · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Cited alongside, same era.
When to use parametric models in reinforcement learning?
Van Hasselt, H. P., Hessel, M., and Aslanides, J · 2019
Cited alongside, same era.
Training larger networks for deep reinforcement learning
Ota, K., Jha, D. K., and Kanezaki, A · 2021
Later among the works it cites.
Fast and data-efficient training of rainbow: an experimental study on atari
Schmidt, D. and Schmied, T · 2021
Later among the works it cites.
Online and offline reinforcement learning by planning with a learned model
Schrittwieser, J., Hubert, T., Mandhane, A., Barekatain, M., Antonoglou, I., and Silver, D · 2021
Later among the works it cites.
Data-efficient reinforcement learning with self-predictive representations
Schwarzer, M., Anand, A., Goel, R., Hjelm, R. D., Courville, A., and Bachman, P · 2021
Later among the works it cites.
Reincarnating reinforcement learning: Reusing prior computation to accelerate progress
Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A., and Bellemare, M. G · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Cited alongside, same era.
An optimistic perspective on offline reinforcement learning
Agarwal, R., Schuurmans, D., and Norouzi, M · 2020
Cited alongside, same era.
Leveraging procedural generation to benchmark reinforcement learning
Cobbe, K., Hesse, C., Hilton, J., and Schulman, J · 2020
Cited alongside, same era.
Bootstrap your own latent-a new approach to self-supervised learning
Grill, J.-B., Strub, F., Altché, F., Tallec, C., Richemond, P., Buchatskaya, E., Doersch, C., Avila Pires, B., Guo, Z., Gheshlaghi Azar, M., et al · 2020
Cited alongside, same era.
Mastering atari with discrete world models
Hafner, D., Lillicrap, T., Norouzi, M., and Ba, J · 2020
Cited alongside, same era.
Array programming with numpy
Harris, C. R., Millman, K. J., Van Der Walt, S. J., Gommers, R., Virtanen, P., Cournapeau, D., Wieser, E., Taylor, J., Berg, S., Smith, N. J., et al · 2020
Cited alongside, same era.
Model-based reinforcement learning for atari
Kaiser, L., Babaeizadeh, M., Milos, P., Osinski, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., et al · 2020
Cited alongside, same era.
Magnetic control of tokamak plasmas through deep reinforcement learning
Degrave, J., Felici, F., Buchli, J., Neunert, M., Tracey, B. D., Carpanese, F., Ewalds, T., Hafner, R., Abdolmaleki, A., de Las Casas, D., Donner, C., Fritz, L., Galperti, C., Huber, A., Keeling, J., Tsimpoukelli, M., Kay, J., Merle, A., Moret, J., Noury, S., Pesamosca, F., Pfau, D., Sauter, O., Sommariva, C., Coda, S., Duval, B., Fasoli, A., Kohli, P., Kavukcuoglu, K., Hassabis, D., and Riedmiller, M. A · 2022
Later among the works it cites.
Minedojo: Building open-ended embodied agents with internet-scale knowledge
Fan, L., Wang, G., Jiang, Y., Mandlekar, A., Yang, Y., Zhu, H., Tang, A., Huang, D.-A., Zhu, Y., and Anandkumar, A · 2022
Later among the works it cites.
Kapturowski, S., Campos, V., Jiang, R., Rakićević, N., van Hasselt, H., Blundell, C., and Badia, A. P · 2022
Later among the works it cites.
Offline q-learning on diverse multi-task data both scales and generalizes, 2022
Kumar, A., Agarwal, R., Geng, X., Tucker, G., and Levine, S · 2022
Later among the works it cites.
Does self-supervised learning really improve reinforcement learning from pixels?
Li, X., Shang, J., Das, S., and Ryoo, M. S · 2022
Later among the works it cites.
The primacy bias in deep reinforcement learning
Nikishin, E., Schwarzer, M., D’Oro, P., Bacon, P.-L., and Courville, A · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Gray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Later among the works it cites.
The phenomenon of policy churn
Schaul, T., Barreto, A., Quan, J., and Ostrovski, G · 2022
Later among the works it cites.
Sample-efficient reinforcement learning by breaking the replay ratio barrier
D’Oro, P., Schwarzer, M., Nikishin, E., Bacon, P.-L., Bellemare, M. G., and Courville, A · 2023
Closest in time.
Proto-value networks: Scaling representation learning with auxiliary tasks
Farebrother, J., Greaves, J., Agarwal, R., Lan, C. L., Goroshin, R., Castro, P. S., and Bellemare, M. G · 2023
Closest in time.
Mastering diverse domains through world models, 2023
Hafner, D., Pasukonis, J., Ba, J., and Lillicrap, T · 2023
Closest in time.
Transformers are sample-efficient world models
Micheli, V., Alonso, E., and Fleuret, F · 2023
Closest in time.
Transformer-based world models are happy with 100k interactions
Robine, J., Höftmann, M., Uelwer, T., and Harmeling, S · 2023
Closest in time.
The dormant neuron phenomenon in deep reinforcement learning
Sokar, G., Agarwal, R., Castro, P. S., and Evci, U · 2023
Closest in time.
Investigating multi-task pretraining and generalization in reinforcement learning
Taiga, A. A., Agarwal, R., Farebrother, J., Courville, A., and Bellemare, M. G · 2023
Closest in time.
Human-timescale adaptation in an open-ended task space
Team, A. A., Bauer, J., Baumli, K., Baveja, S., Behbahani, F., Bhoopchand, A., Bradley-Schmieg, N., Chang, M., Clay, N., Collister, A., et al · 2023
Closest in time.