Fetching the paper…
Reading the bibliography…
Recent work has shown that deep reinforcement learning agents have difficulty in effectively using their network parameters.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Python reference manual
Van Rossum, G. and Drake Jr, F. L · 1995
Earlier work this paper cites.
Matplotlib: A 2d graphics environment
Hunter, J. D · 2007
Earlier work this paper cites.
Python for scientific computing
Oliphant, T. E · 2007
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Python for Data Analysis: Data Wrangling with Pandas, NumPy, and IPython
McKinney, W · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Han, S., Mao, H., and Dally, W. J · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z · 2016
Earlier work this paper cites.
Jupyter Notebooks—a publishing format for reproducible computational workflows
Kluyver, T., Ragan-Kelley, B., Pérez, F., Granger, B., Bussonnier, M., Frederic, J., Kelley, K., Hamrick, J., Grout, J., Corlay, S., Ivanov, P., Avila, D., Abdalla, S., Willing, C., and Jupyter Development Team · 2016
Earlier work this paper cites.
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
van Hasselt, H., Guez, A., and Silver, D · 2016
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Wang, Z., Schaul, T., Hessel, M., Hasselt, H., Lanctot, M., and Freitas, N · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Bellemare, M. G., Dabney, W., and Munos, R · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
To prune, or not to prune: exploring the efficacy of pruning for model compression, 2017
Zhu, M. and Gupta, S · 2017
Earlier work this paper cites.
Jax: composable transformations of python+ numpy programs
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., et al · 2018
Earlier work this paper cites.
Dopamine: A research framework for deep reinforcement learning
Castro, P. S., Moitra, S., Gelada, C., Kumar, S., and Bellemare, M. G · 2018
Earlier work this paper cites.
Implicit quantile networks for distributional reinforcement learning
Dabney, W., Ostrovski, G., Silver, D., and Munos, R · 2018
Earlier work this paper cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., et al · 2018
Earlier work this paper cites.
Noisy networks for exploration
Fortunato, M., Azar, M. G., Piot, B., Menick, J., Osband, I., Graves, A., Mnih, V., Munos, R., Hassabis, D., Pietquin, O., Blundell, C., and Legg, S · 2018
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., Hasselt, H. V., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M. G., and Silver, D · 2018
Earlier work this paper cites.
Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science
Mocanu, D. C., Mocanu, E., Stone, P., Nguyen, P. H., Gibescu, M., and Liotta, A · 2018
Earlier work this paper cites.
Kickstarting deep reinforcement learning
Schmitt, S., Hudson, J. J., Zidek, A., Osindero, S., Doersch, C., Czarnecki, W. M., Leibo, J. Z., Kuttler, H., Zisserman, A., Simonyan, K., et al · 2018
Cited alongside, same era.
Deep reinforcement learning and the deadly triad
Van Hasselt, H., Doron, Y., Strub, F., Hessel, M., Sonnerat, N., and Modayil, J · 2018
Cited alongside, same era.
A study on overfitting in deep reinforcement learning
Zhang, C., Vinyals, O., Munos, R., and Bengio, S · 2018
Cited alongside, same era.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., et al · 2019
Cited alongside, same era.
Training larger networks for deep reinforcement learning
Ota, K., Jha, D. K., and Kanezaki, A · 2021
Later among the works it cites.
Fast and data-efficient training of rainbow: an experimental study on atari
Schmidt, D. and Schmied, T · 2021
Later among the works it cites.
Dynamic sparse training for deep reinforcement learning
Sokar, G., Mocanu, E., Mocanu, D. C., Pechenizkiy, M., and Stone, P · 2021
Later among the works it cites.
On lottery tickets and minimal task representations in deep reinforcement learning
Vischer, M., Lange, R. T., and Sprekeler, H · 2021
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Yarats, D., Fergus, R., and Kostrikov, I · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gale, T., Elsen, E., and Hooker, S · 2019
Cited alongside, same era.
Observational overfitting in reinforcement learning
Song, X., Jiang, Y., Tu, S., Du, Y., and Neyshabur, B · 2019
Cited alongside, same era.
When to use parametric models in reinforcement learning?
Van Hasselt, H. P., Hessel, M., and Aslanides, J · 2019
Cited alongside, same era.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Cited alongside, same era.
Playing the lottery with rewards and multiple languages: lottery tickets in rl and nlp
Yu, H., Edunov, S., Tian, Y., and Morcos, A. S · 2019
Cited alongside, same era.
Accelerating the deep reinforcement learning with neural network compression
Zhang, H., He, Z., and Li, J · 2019
Cited alongside, same era.
An optimistic perspective on offline reinforcement learning
Agarwal, R., Schuurmans, D., and Norouzi, M · 2020
Cited alongside, same era.
Autonomous navigation of stratospheric balloons using reinforcement learning
Bellemare, M. G., Candido, S., Castro, P. S., Gong, J., Machado, M. C., Moitra, S., Ponda, S. S., and Wang, Z · 2020
Cited alongside, same era.
Later among the works it cites.
Reincarnating reinforcement learning: Reusing prior computation to accelerate progress
Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A. C., and Bellemare, M · 2022
Later among the works it cites.
Stabilizing off-policy deep reinforcement learning from pixels
Cetin, E., Ball, P. J., Roberts, S., and Celiktutan, O · 2022
Later among the works it cites.
Discovering faster matrix multiplication algorithms with reinforcement learning
Fawzi, A., Balog, M., Huang, A., Hubert, T., Romera-Paredes, B., Barekatain, M., Novikov, A., R Ruiz, F. J., Schrittwieser, J., Swirszcz, G., et al · 2022
Later among the works it cites.
The state of sparse training in deep reinforcement learning
Graesser, L., Evci, U., Elsen, E., and Castro, P. S · 2022
Later among the works it cites.
Offline q-learning on diverse multi-task data both scales and generalizes
Kumar, A., Agarwal, R., Geng, X., Tucker, G., and Levine, S · 2022
Later among the works it cites.
Understanding and preventing capacity loss in reinforcement learning
Lyle, C., Rowland, M., and Dabney, W · 2022
Later among the works it cites.
The primacy bias in deep reinforcement learning
Nikishin, E., Schwarzer, M., D’Oro, P., Bacon, P.-L., and Courville, A · 2022
Later among the works it cites.
Small batch deep reinforcement learning
Ceron, J. S. O., Bellemare, M. G., and Castro, P. S · 2023
Later among the works it cites.
Sample-efficient reinforcement learning by breaking the replay ratio barrier
D’Oro, P., Schwarzer, M., Nikishin, E., Bacon, P.-L., Bellemare, M. G., and Courville, A · 2023
Later among the works it cites.
Proto-value networks: Scaling representation learning with auxiliary tasks
Farebrother, J., Greaves, J., Agarwal, R., Lan, C. L., Goroshin, R., Castro, P. S., and Bellemare, M. G · 2023
Later among the works it cites.
Automatic noise filtering with dynamic sparse training in deep reinforcement learning
Grooten, B., Sokar, G., Dohare, S., Mocanu, E., Taylor, M. E., Pechenizkiy, M., and Mocanu, D. C · 2023
Later among the works it cites.
Mastering diverse domains through world models
Hafner, D., Pasukonis, J., Ba, J., and Lillicrap, T · 2023
Later among the works it cites.
Plastic: Improving input and label plasticity for sample efficient reinforcement learning
Lee, H., Cho, H., Kim, H., Gwak, D., Kim, J., Choo, J., Yun, S.-Y., and Yun, C · 2023
Later among the works it cites.
Curvature explains loss of plasticity
Lewandowski, A., Tanaka, H., Schuurmans, D., and Machado, M. C · 2023
Later among the works it cites.
Understanding plasticity in neural networks
Lyle, C., Zheng, Z., Nikishin, E., Pires, B. A., Pascanu, R., and Dabney, W · 2023
Later among the works it cites.
Deep reinforcement learning with plasticity injection
Nikishin, E., Oh, J., Ostrovski, G., Lyle, C., Pascanu, R., Dabney, W., and Barreto, A · 2023
Later among the works it cites.
Bigger, better, faster: Human-level atari with human-level efficiency
Schwarzer, M., Ceron, J. S. O., Courville, A., Bellemare, M. G., Agarwal, R., and Castro, P. S · 2023
Later among the works it cites.
The dormant neuron phenomenon in deep reinforcement learning
Sokar, G., Agarwal, R., Castro, P. S., and Evci, U · 2023
Later among the works it cites.
Investigating multi-task pretraining and generalization in reinforcement learning
Taiga, A. A., Agarwal, R., Farebrother, J., Courville, A., and Bellemare, M. G · 2023
Later among the works it cites.
RLx2: Training a sparse deep reinforcement learning model from scratch
Tan, Y., Hu, P., Pan, L., Huang, J., and Huang, L · 2023
Later among the works it cites.
Mixtures of experts unlock parameter scaling for deep RL
Ceron, J. S. O., Sokar, G., Willi, T., Lyle, C., Farebrother, J., Foerster, J. N., Dziugaite, G. K., Precup, D., and Castro, P. S · 2024
Closest in time.
Stop regressing: Training value functions via classification for scalable deep rl
Farebrother, J., Orbay, J., Vuong, Q., Taïga, A. A., Chebotar, Y., Xiao, T., Irpan, A., Levine, S., Castro, P. S., Faust, A., Kumar, A., and Agarwal, R · 2024
Closest in time.
Jaxpruner: A concise library for sparsity research
Lee, J. H., Park, W., Mitchell, N. E., Pilault, J., Ceron, J. S. O., Kim, H.-B., Lee, N., Frantar, E., Long, Y., Yazdanbakhsh, A., et al · 2024
Closest in time.
Disentangling the causes of plasticity loss in neural networks
Lyle, C., Zheng, Z., Khetarpal, K., van Hasselt, H., Pascanu, R., Martens, J., and Dabney, W · 2024
Closest in time.