Fetching the paper…
Reading the bibliography…
The Mixture-of-Expert (MoE) technique plays a crucial role in expanding the size of DNN model parameters.
Knapsack problems: algorithms and computer implementations
Martello, S. and Toth, P · 1990
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q. V., Hinton, G. E., and Dean, J · 2017
Earlier work this paper cites.
TVM: An automated end-to-end optimizing compiler for deep learning
Chen, T., Moreau, T., Jiang, Z., Zheng, L., Yan, E., Shen, H., Cowan, M., Wang, L., Hu, Y., Ceze, L., Guestrin, C., and Krishnamurthy, A · 2018
Earlier work this paper cites.
Glow: Graph lowering compiler techniques for neural networks
Rotem, N., Fix, J., Abdulrasool, S., Catron, G., Deng, S., Dzhabarov, R., Gibson, N., Hegeman, J., Lele, M., Levenstein, R., et al · 2018
Earlier work this paper cites.
Priority-based parameter propagation for distributed dnn training
Jayarajan, A., Jinliang, W., Gibson, G., Fedorova, A., and Pekhimenko, G · 2019
Earlier work this paper cites.
PyTorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Earlier work this paper cites.
A generic communication scheduler for distributed dnn training acceleration
Peng, Y., Zhu, Y., Chen, Y., Bao, Y., Yi, B., Lan, C., Wu, C., and Guo, C · 2019
Earlier work this paper cites.
OR-Tools
Perron, L. and Furnon, V · 2019
Earlier work this paper cites.
Language Models are Unsupervised Multitask Learners, 2019
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 2019
Earlier work this paper cites.
GShard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z · 2020
Earlier work this paper cites.
ZeRO: Memory optimizations toward training trillion parameter models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y · 2020
Earlier work this paper cites.
DeepSpeed: System optimizations enable training deep learning models with over 100 billion parameters
Rasley, J., Rajbhandari, S., Ruwase, O., and He, Y · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M · 2020
Cited alongside, same era.
M6-10t: A sharing-delinking paradigm for efficient multi-trillion parameter pretraining
Lin, J., Yang, A., Bai, J., Zhou, C., Jiang, L., Jia, X., Wang, A., Zhang, J., Li, Y., Lin, W., et al · 2021
Cited alongside, same era.
MPI: A Message-Passing Interface Standard Version 4.0 , June 2021
Message Passing Interface Forum · 2021
Cited alongside, same era.
NCCL, 2021
NVIDIA · 2021
Cited alongside, same era.
Scaling vision with sparse mixture of experts
Overlap communication with dependent computation via decomposition in large deep learning models
Wang, S., Wei, J., Sabne, A., Davis, A., Ilbeyi, B., Hechtman, B., Chen, D., Murthy, K. S., Maggioni, M., Zhang, Q., et al · 2022
Later among the works it cites.
Accelerating large-scale distributed neural network training with SPMD parallelism
Zhang, S., Diao, L., Wu, C., Wang, S., and Lin, W · 2022
Later among the works it cites.
Alpa: Automating inter- and intra-operator parallelism for distributed deep learning
Zheng, L., Li, Z., Zhang, H., Zhuang, Y., Chen, Z., Huang, Y., Wang, Y., Xu, Y., Zhuo, D., Gonzalez, J. E., et al · 2022
Later among the works it cites.
Mixture-of-experts with expert choice routing
Zhou, Y., Lei, T., Liu, H., Du, N., Huang, Y., Zhao, V. Y., Dai, A. M., Chen, Z., Le, Q. V., and Laudon, J · 2022
Later among the works it cites.
Taming sparsely activated transformer with stochastic experts
Zuo, S., Liu, X., Jiao, J., Kim, Y. J., Hassan, H., Zhang, R., Gao, J., and Zhao, T · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Riquelme, C., Puigcerver, J., Mustafa, B., Neumann, M., Jenatton, R., Susano Pinto, A., Keysers, D., and Houlsby, N · 2021
Cited alongside, same era.
Hash layers for large sparse models
Roller, S., Sukhbaatar, S., Weston, J., et al · 2021
Cited alongside, same era.
M6-t: Exploring sparse expert models and beyond
Yang, A., Lin, J., Men, R., Zhou, C., Jiang, L., Jia, X., Wang, A., Zhang, J., Wang, J., Li, Y., et al · 2021
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2022
Cited alongside, same era.
FasterMoE: modeling and optimizing training of large-scale dynamic pre-trained models
He, J., Zhai, J., Antunes, T., Wang, H., Luo, F., Shi, S., and Li, Q · 2022
Cited alongside, same era.
HetuMoE: An efficient trillion-scale mixture-of-expert distributed training system
Nie, X., Zhao, P., Miao, X., and Cui, B · 2022
Cited alongside, same era.
DeepSpeed-MoE: Advancing mixture-of-experts inference and training to power next-generation AI scale
Rajbhandari, S., Li, C., Yao, Z., Zhang, M., Aminabadi, R. Y., Awan, A. A., Rasley, J., and He, Y · 2022
Cited alongside, same era.
Sparse moe as the new dropout: Scaling dense and self-slimmable transformers
Chen, T., Zhang, Z., Jaiswal, A. K., Liu, S., and Wang, Z · 2023
Later among the works it cites.
MegaBlocks: Efficient Sparse Training with Mixture-of-Experts
Gale, T., Narayanan, D., Young, C., and Zaharia, M · 2023
Later among the works it cites.
Tutel: Adaptive mixture-of-experts at scale
Hwang, C., Cui, W., Xiong, Y., Yang, Z., Liu, Z., Hu, H., Wang, Z., Salas, R., Jose, J., Ram, P., et al · 2023
Later among the works it cites.
Ring attention with blockwise transformers for near-infinite context
Liu, H., Zaharia, M., and Abbeel, P · 2023
Later among the works it cites.
RAF: Holistic compilation for deep learning model training
Yu, C. H., Fan, H., Huang, G., Jia, Z., Liu, Y., Wang, J., Zheng, Z., Zhou, Y., Shen, H., Shao, J., et al · 2023
Later among the works it cites.
DeepSeekMoE: Towards ultimate expert specialization in mixture-of-experts language models
Dai, D., Deng, C., Zhao, C., Xu, R. X., Gao, H., Chen, D., Li, J., Zeng, W., Yu, X., Wu, Y., Xie, Z., Li, Y. K., Huang, P., Luo, F., Ruan, C., Sui, Z., and Liang, W · 2024
Closest in time.
Mixtral of experts
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., de las Casas, D., Hanna, E. B., Bressand, F., Lengyel, G., Bour, G., Lample, G., Lavaud, L. R., Saulnier, L., Lachaux, M.-A., Stock, P., Subramanian, S., Yang, S., Antoniak, S., Scao, T. L., Gervet, T., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2024
Closest in time.