Fetching the paper…
Reading the bibliography…
The Mixture of Experts architecture allows for outrageously large neural networks by scaling model parameter size independently from computational demand (FLOPs).
Backpropagation applied to handwritten zip code recognition
LeCun, Y., Boser, B., Denker, J. S., Henderson, D., Howard, R. E., Hubbard, W., and Jackel, L. D · 1989
Earlier work this paper cites.
Adaptive mixture of local expert
Jacobs, R., Jordan, M., Nowlan, S., and Hinton, G · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the em algorithm
Jordan, M. and Jacobs, R · 1993
Earlier work this paper cites.
Legion: Expressing locality and independence with logical regions
Bauer, M., Treichler, S., Slaughter, E., and Aiken, A · 2012
Earlier work this paper cites.
Caffe: Convolutional architecture for fast feature embedding
Jia, Y., Shelhamer, E., Donahue, J., Karayev, S., Long, J., Girshick, R., Guadarrama, S., and Darrell, T · 2014
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X · 2015
Earlier work this paper cites.
Cntk: Microsoft’s open-source deep-learning toolkit
Seide, F. and Agarwal, A · 2016
Earlier work this paper cites.
Theano: A Python framework for fast computation of mathematical expressions
Theano Development Team · 2016
Earlier work this paper cites.
Dynet: The dynamic neural network toolkit
Neubig, G., Dyer, C., Goldberg, Y., Matthews, A., Ammar, W., Anastasopoulos, A., Ballesteros, M., Chiang, D., Clothiaux, D., Cohn, T., Duh, K., Faruqui, M., Gan, C., Garrette, D., Ji, Y., Kong, L., Kuncoro, A., Kumar, G., Malaviya, C., Michel, P., Oda, Y., Richardson, M., Saphra, N., Swayamdipta, S., and Yin, P · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer, 2017
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J · 2017
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., and Zhang, Q · 2018
Cited alongside, same era.
Beyond data and model parallelism for deep neural networks, 2018
Jia, Z., Zaharia, M., and Aiken, A · 2018
Cited alongside, same era.
Tensor2Tensor for neural machine translation
Vaswani, A., Bengio, S., Brevdo, E., Chollet, F., Gomez, A., Gouws, S., Jones, L., Kaiser, Ł., Kalchbrenner, N., Parmar, N., Sepassi, R., Shazeer, N., and Uszkoreit, J · 2018
Cited alongside, same era.
Chainer: A deep learning framework for accelerating the research cycle
Tokui, S., Okuta, R., Akiba, T., Niitani, Y., Ogawa, T., Saito, S., Suzuki, S., Uenishi, K., Vogel, B., and Yamazaki Vincent, H · 2019
Cited alongside, same era.
Fastmoe: A fast mixture-of-expert training system, 2021
He, J., Qiu, J., Zeng, A., Yang, Z., Zhai, J., and Tang, J · 2021
Later among the works it cites.
The CIFAR10 dataset, 2014
Krizhevsky, A., Nair, V., and Hinton, G · 2021
Later among the works it cites.
Sparsely-gated mixture-of-experts pytorch implementation, 2019
Rau, D · 2021
Later among the works it cites.
Scaling vision with sparse mixture of experts, 2021
Riquelme, C., Puigcerver, J., Mustafa, B., Neumann, M., Jenatton, R., Pinto, A. S., Keysers, D., and Houlsby, N · 2021
Later among the works it cites.
The bitter lesson, 2019
Sutton, R · 2021
Later among the works it cites.
Sparsely-gated mixture-of-experts pytorch implementation, 2019
Wang, P · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Language models are few-shot learners, 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Scaling laws for neural language models, 2020
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Cited alongside, same era.
Gshard: Scaling giant models with conditional computation and automatic sharding, 2020
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z · 2020
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2021
Cited alongside, same era.
Keras callbacks with tensorflow
TensorFlow
Cited in the paper.
ST-MoE: Designing Stable and Transferable Sparse Expert Models, 2022
Zoph, B., Bello, I., Kumar, S., Du, N., Huang, Y., Dean, J., Shazeer, N., and Fedus, W · 2022
Closest in time.
Superglue: A stickier benchmark for general-purpose language understanding systems, 2020
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2022
Closest in time.