Fetching the paper…
Reading the bibliography…
The recent rapid progress in (self) supervised learning models is in large part predicted by empirical scaling laws: a model's performance scales proportionally to its size.
The state of sparsity in deep neural networks
Gale, T., Elsen, E., and Hooker, S · 1902
Earlier work this paper cites.
Visual feature extraction by a multilayered network of analog threshold elements
Fukushima, K · 1969
Earlier work this paper cites.
Adaptive mixtures of local experts
Jacobs, R. A., Jordan, M. I., Nowlan, S. J., and Hinton, G. E · 1991
Earlier work this paper cites.
Python reference manual
Van Rossum, G. and Drake Jr, F. L · 1995
Earlier work this paper cites.
Introduction to Reinforcement Learning
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Matplotlib: A 2d graphics environment
Hunter, J. D · 2007
Earlier work this paper cites.
Python for scientific computing
Oliphant, T. E · 2007
Earlier work this paper cites.
Python for Data Analysis: Data Wrangling with Pandas, NumPy, and IPython
McKinney, W · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., and Hassabis, D · 2015
Earlier work this paper cites.
Jupyter Notebooks—a publishing format for reproducible computational workflows
Kluyver, T., Ragan-Kelley, B., Pérez, F., Granger, B., Bussonnier, M., Frederic, J., Kelley, K., Hamrick, J., Grout, J., Corlay, S., Ivanov, P., Avila, D., Abdalla, S., Willing, C., and Jupyter Development Team · 2016
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
Jax: composable transformations of python+ numpy programs
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., et al · 2018
Earlier work this paper cites.
Dopamine: A Research Framework for Deep Reinforcement Learning
Castro, P. S., Moitra, S., Gelada, C., Kumar, S., and Bellemare, M. G · 2018
Earlier work this paper cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., et al · 2018
Earlier work this paper cites.
Rainbow: Combining improvements in deep reinforcement learning
Hessel, M., Modayil, J., Hasselt, H. V., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M. G., and Silver, D · 2018
Earlier work this paper cites.
Revisiting the arcade learning environment: evaluation protocols and open problems for general agents
Machado, M. C., Bellemare, M. G., Talvitie, E., Veness, J., Hausknecht, M., and Bowling, M · 2018
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., et al · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S · 2019
Earlier work this paper cites.
A neural dirichlet process mixture model for task-free continual learning
Lee, S., Ha, J., Zhang, D., and Kim, G · 2019
Earlier work this paper cites.
When to use parametric models in reinforcement learning?
Van Hasselt, H. P., Hessel, M., and Aslanides, J · 2019
Earlier work this paper cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Earlier work this paper cites.
Condconv: Conditionally parameterized convolutions for efficient inference
Yang, B., Bender, G., Le, Q. V., and Ngiam, J · 2019
Earlier work this paper cites.
Biased mixtures of experts: Enabling computer vision inference under data transfer limitations
Abbas, A. and Andreopoulos, Y · 2020
Earlier work this paper cites.
An optimistic perspective on offline reinforcement learning
Agarwal, R., Schuurmans, D., and Norouzi, M · 2020
Earlier work this paper cites.
Autonomous navigation of stratospheric balloons using reinforcement learning
Bellemare, M. G., Candido, S., Castro, P. S., Gong, J., Machado, M. C., Moitra, S., Ponda, S. S., and Wang, Z · 2020
Earlier work this paper cites.
Rigging the lottery: Making all tickets winners
Evci, U., Gale, T., Menick, J., Castro, P. S., and Elsen, E · 2020
Earlier work this paper cites.
Revisiting fundamentals of experience replay
Fedus, W., Ramachandran, P., Agarwal, R., Bengio, Y., Larochelle, H., Rowland, M., and Dabney, W · 2020
Cited alongside, same era.
Array programming with numpy
Harris, C. R., Millman, K. J., Van Der Walt, S. J., Gommers, R., Virtanen, P., Cournapeau, D., Wieser, E., Taylor, J., Berg, S., Smith, N. J., et al · 2020
Cited alongside, same era.
Transient non-stationarity and generalisation in deep reinforcement learning
Igl, M., Farquhar, G., Luketina, J., Boehmer, W., and Whiteson, S · 2020
Cited alongside, same era.
Model based reinforcement learning for atari
Kaiser, L., Babaeizadeh, M., Miłos, P., Osiński, B., Campbell, R. H., Czechowski, K., Erhan, D., Finn, C., Kozakowski, P., Levine, S., Mohiuddin, A., Sepassi, R., Tucker, G., and Michalewski, H · 2020
Cited alongside, same era.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Cited alongside, same era.
M 3 vit: Mixture-of-experts vision transformer for efficient multi-task learning with model-accelerator co-design
Fan, Z., Sarkar, R., Jiang, Z., Chen, T., Zou, K., Cheng, Y., Hao, C., Wang, Z., et al · 2022
Later among the works it cites.
Proto-value networks: Scaling representation learning with auxiliary tasks
Farebrother, J., Greaves, J., Agarwal, R., Le Lan, C., Goroshin, R., Castro, P. S., and Bellemare, M. G · 2022
Later among the works it cites.
Discovering faster matrix multiplication algorithms with reinforcement learning
Fawzi, A., Balog, M., Huang, A., Hubert, T., Romera-Paredes, B., Barekatain, M., Novikov, A., R Ruiz, F. J., Schrittwieser, J., Swirszcz, G., et al · 2022
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2022
Later among the works it cites.
The state of sparse training in deep reinforcement learning
Graesser, L., Evci, U., Elsen, E., and Castro, P. S · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Gshard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z · 2020
Cited alongside, same era.
Using mixture of expert models to gain insights into semantic segmentation
Pavlitskaya, S., Hubschneider, C., Weber, M., Moritz, R., Huger, F., Schlicht, P., and Zollner, M · 2020
Cited alongside, same era.
Scalable transfer learning with expert models
Puigcerver, J., Ruiz, C. R., Mustafa, B., Renggli, C., Pinto, A. S., Gelly, S., Keysers, D., and Houlsby, N · 2020
Cited alongside, same era.
Deep mixture of experts via shallow embedding
Wang, X., Yu, F., Dunlap, L., Ma, Y.-A., Wang, R., Mirhoseini, A., Darrell, T., and Gonzalez, J. E · 2020
Cited alongside, same era.
Deep reinforcement learning at the edge of the statistical precipice
Agarwal, R., Schwarzer, M., Castro, P. S., Courville, A. C., and Bellemare, M · 2021
Cited alongside, same era.
Continuous action reinforcement learning from a mixture of interpretable experts
Akrour, R., Tateo, D., and Peters, J · 2021
Cited alongside, same era.
An empirical study of implicit regularization in deep offline rl
Gulcehre, C., Srinivasan, S., Sygnowski, J., Ostrovski, G., Farajtabar, M., Hoffman, M., Pascanu, R., and Doucet, A · 2022
Later among the works it cites.
Offline q-learning on diverse multi-task data both scales and generalizes
Kumar, A., Agarwal, R., Geng, X., Tucker, G., and Levine, S · 2022
Later among the works it cites.
Multimodal contrastive learning with limoe: the language-image mixture of experts
Mustafa, B., Riquelme, C., Puigcerver, J., Jenatton, R., and Houlsby, N · 2022
Later among the works it cites.
The primacy bias in deep reinforcement learning
Nikishin, E., Schwarzer, M., D’Oro, P., Bacon, P.-L., and Courville, A · 2022
Later among the works it cites.
Dynamic sparse training for deep reinforcement learning
Sokar, G., Mocanu, E., Mocanu, D. C., Pechenizkiy, M., and Stone, P · 2022
Later among the works it cites.
Investigating multi-task pretraining and generalization in reinforcement learning
Taiga, A. A., Agarwal, R., Farebrother, J., Courville, A., and Bellemare, M. G · 2022
Later among the works it cites.
Rlx2: Training a sparse deep reinforcement learning model from scratch
Tan, Y., Hu, P., Pan, L., Huang, J., and Huang, L · 2022
Later among the works it cites.
Mixture-of-experts with expert choice routing
Zhou, Y., Lei, T., Liu, H., Du, N., Huang, Y., Zhao, V., Dai, A. M., Le, Q. V., Laudon, J., et al · 2022
Later among the works it cites.
St-moe: Designing stable and transferable sparse expert models, 2022
Zoph, B., Bello, I., Kumar, S., Du, N., Huang, Y., Dean, J., Shazeer, N., and Fedus, W · 2022
Later among the works it cites.
Small batch deep reinforcement learning
Ceron, J. S. O., Bellemare, M. G., and Castro, P. S · 2023
Later among the works it cites.
Mod-squad: Designing mixtures of experts as modular multi-task learners
Chen, Z., Shen, Y., Ding, M., Chen, Z., Zhao, H., Learned-Miller, E. G., and Gan, C · 2023
Later among the works it cites.
Sample-efficient reinforcement learning by breaking the replay ratio barrier
D’Oro, P., Schwarzer, M., Nikishin, E., Bacon, P.-L., Bellemare, M. G., and Courville, A · 2023
Later among the works it cites.
MegaBlocks: Efficient Sparse Training with Mixture-of-Experts
Gale, T., Narayanan, D., Young, C., and Zaharia, M · 2023
Later among the works it cites.
Relu to the rescue: Improve your on-policy actor-critic with positive advantages
Jesson, A., Lu, C., Gupta, G., Filos, A., Foerster, J. N., and Gal, Y · 2023
Later among the works it cites.
From sparse to soft mixtures of experts, 2023
Puigcerver, J., Riquelme, C., Mustafa, B., and Houlsby, N · 2023
Later among the works it cites.
Bigger, better, faster: Human-level Atari with human-level efficiency
Schwarzer, M., Obando Ceron, J. S., Courville, A., Bellemare, M. G., Agarwal, R., and Castro, P. S · 2023
Later among the works it cites.
The dormant neuron phenomenon in deep reinforcement learning
Sokar, G., Agarwal, R., Castro, P. S., and Evci, U · 2023
Later among the works it cites.
Taskexpert: Dynamically assembling multi-task representations with memorial mixture-of-experts
Ye, H. and Xu, D · 2023
Later among the works it cites.
In value-based deep reinforcement learning, a pruned network is a good network
Ceron, J. S. O., Courville, A., and Castro, P. S · 2024
Closest in time.
Stop regressing: Training value functions via classification for scalable deep rl
Farebrother, J., Orbay, J., Vuong, Q., Taïga, A. A., Chebotar, Y., Xiao, T., Irpan, A., Levine, S., Castro, P. S., Faust, A., Kumar, A., and Agarwal, R · 2024
Closest in time.
Multi-task reinforcement learning with mixture of orthogonal experts
Hendawy, A., Peters, J., and D’Eramo, C · 2024
Closest in time.