Fetching the paper…
Reading the bibliography…
Mixture-of-Experts (MoE) has emerged as a practical approach to scale up parameters for the Transformer model to achieve better generalization while maintaining a sub-linear increase in computation overhead.
Attention is all you need
Vaswani, A · 2017
Earlier work this paper cites.
Volta: Performance and programmability
Choquette, J., Giroux, O., and Foley, D · 2018
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Fastmoe: A fast mixture-of-expert training system
He, J., Qiu, J., Zeng, A., Yang, Z., Zhai, J., and Tang, J · 2021
Earlier work this paper cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B · 2021
Earlier work this paper cites.
Efficient large-scale language model training on gpu clusters using megatron-lm
Narayanan, D., Shoeybi, M., Casper, J., LeGresley, P., Patwary, M., Korthikanti, V., Vainbrand, D., Kashinkunti, P., Bernauer, J., Catanzaro, B., et al · 2021
Earlier work this paper cites.
Scaling vision with sparse mixture of experts
Riquelme, C., Puigcerver, J., Mustafa, B., Neumann, M., Jenatton, R., Susano Pinto, A., Keysers, D., and Houlsby, N · 2021
Cited alongside, same era.
Vlmo: Unified vision-language pre-training with mixture-of-modality-experts
Bao, H., Wang, W., Dong, L., Liu, Q., Mohammed, O. K., Aggarwal, K., Som, S., Piao, S., and Wei, F · 2022
Cited alongside, same era.
Glam: Efficient scaling of language models with mixture-of-experts
Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., et al · 2022
Cited alongside, same era.
M 3 vit: Mixture-of-experts vision transformer for efficient multi-task learning with model-accelerator co-design
Fan, Z., Sarkar, R., Jiang, Z., Chen, T., Zou, K., Cheng, Y., Hao, C., Wang, Z., et al · 2022
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2022
Cited alongside, same era.
Accelerating distributed { \{ MoE } \} training and inference with lina
Li, J., Jiang, Y., Zhu, Y., Wang, C., and Xu, H · 2023
Later among the works it cites.
Janus: A unified distributed training framework for sparse mixture-of-experts models
Liu, J., Wang, J. H., and Jiang, Y · 2023
Later among the works it cites.
Pipemoe: Accelerating mixture-of-experts through adaptive pipelining
Shi, S., Pan, X., Chu, X., and Li, B · 2023
Later among the works it cites.
{ \{ SmartMoE } \} : Efficiently training { \{ Sparsely-Activated } \} models through combining offline and online parallelization
Zhai, M., He, J., Ma, Z., Zong, Z., Zhang, R., and Zhai, J · 2023
Later among the works it cites.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J., Li, D., Xiong, C., and Hoi, S · 2022
Cited alongside, same era.
Megablocks: Efficient sparse training with mixture-of-experts
Gale, T., Narayanan, D., Young, C., and Zaharia, M · 2023
Cited alongside, same era.
Tutel: Adaptive mixture-of-experts at scale
Hwang, C., Cui, W., Xiong, Y., Yang, Z., Liu, Z., Hu, H., Wang, Z., Salas, R., Jose, J., Ram, P., et al · 2023
Cited alongside, same era.
Fastermoe: modeling and optimizing training of large-scale dynamic pre-trained models
He, J., Zhai, J., Antunes, T., Wang, H., Luo, F., Shi, S., and Li, Q
Cited in the paper.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R
Cited in the paper.
Visual instruction tuning
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2024
Closest in time.
Schemoe: An extensible mixture-of-experts distributed training system with tasks scheduling
Shi, S., Pan, X., Wang, Q., Liu, C., Ren, X., Hu, Z., Yang, Y., Li, B., and Chu, X · 2024
Closest in time.
Mpmoe: Memory efficient moe for pre-trained models with adaptive pipeline parallelism
Zhang, Z., Xia, Y., Wang, H., Yang, D., Hu, C., Zhou, X., and Cheng, D · 2024
Closest in time.