Fetching the paper…
Reading the bibliography…
Expert parallelism has emerged as a key strategy for distributing the computational workload of sparsely-gated mixture-of-experts (MoE) models across multiple devices, enabling the processing of increasingly large-scale models.
Ultra-Performance Pascal GPU and NVLink Interconnect
Foley, D. and Danskin, J · 2017
Earlier work this paper cites.
Densely Connected Convolutional Networks
Huang, G., Liu, Z., van der Maaten, L., and Weinberger, K. Q · 2017
Earlier work this paper cites.
RACE: Large-scale ReAding Comprehension Dataset From Examinations
Lai, G., Xie, Q., Liu, H., Yang, Y., and Hovy, E · 2017
Earlier work this paper cites.
Pointer Sentinel Mixture Models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2017
Earlier work this paper cites.
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q. V., Hinton, G. E., and Dean, J · 2017
Earlier work this paper cites.
Attention is All you Need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Universal Transformers
Dehghani, M., Gouws, S., Vinyals, O., Uszkoreit, J., and Kaiser, L · 2018
Earlier work this paper cites.
Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering
Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A · 2018
Earlier work this paper cites.
MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms
Amini, A., Gabriel, S., Lin, S., Koncel-Kedziorski, R., Choi, Y., and Hajishirzi, H · 2019
Earlier work this paper cites.
BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Earlier work this paper cites.
OpenWebText Corpus
Gokaslan, A. and Cohen, V · 2019
Earlier work this paper cites.
GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism
Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, D., Chen, M. X., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., and Chen, Z · 2019
Earlier work this paper cites.
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R · 2019
Earlier work this paper cites.
ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Lu, J., Batra, D., Parikh, D., and Lee, S · 2019
Earlier work this paper cites.
PipeDream: Generalized Pipeline Parallelism for DNN Training
Narayanan, D., Harlap, A., Phanishayee, A., Seshadri, V., Devanur, N. R., Ganger, G. R., Gibbons, P. B., and Zaharia, M · 2019
Earlier work this paper cites.
Fairseq: A Fast, Extensible Toolkit for Sequence Modeling
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M · 2019
Earlier work this paper cites.
Language Models Are Unsupervised Multitask Learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
HellaSwag: Can a Machine Really Finish Your Sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
PIQA: Reasoning about Physical Commonsense in Natural Language
Bisk, Y., Zellers, R., Gao, J., Choi, Y., et al · 2020
Earlier work this paper cites.
Language Models Are Few-Shot Learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Evaluating Modern GPU Interconnect: PCIe, NVLink, NV-SLI, NVSwitch and GPUDirect
Li, A., Song, S. L., Chen, J., Li, J., Liu, X., Tallent, N. R., and Barker, K. J · 2020
Earlier work this paper cites.
Scalable Deep Learning on Distributed Infrastructures: Challenges, Techniques, and Tools
Mayer, R. and Jacobsen, H.-A · 2020
Earlier work this paper cites.
ZeRO: Memory Optimizations toward Training Trillion Parameter Models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y · 2020
Earlier work this paper cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Cited alongside, same era.
FastMoE: A Fast Mixture-of-Expert Training System
He, J., Qiu, J., Zeng, A., Yang, Z., Zhai, J., and Tang, J · 2021
Cited alongside, same era.
GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z · 2021
Cited alongside, same era.
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B · 2021
Cited alongside, same era.
Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM
Narayanan, D., Shoeybi, M., Casper, J., LeGresley, P., Patwary, M., Korthikanti, V., Vainbrand, D., Kashinkunti, P., Bernauer, J., Catanzaro, B., Phanishayee, A., and Zaharia, M · 2021
Mixture of Cluster-Conditional LoRA Experts for Vision-Language Instruction Tuning
Gou, Y., Liu, Z., Chen, K., Hong, L., Xu, H., Li, A., Yeung, D., Kwok, J. T., and Zhang, Y · 2023
Later among the works it cites.
Tutel: Adaptive Mixture-of-Experts at Scale
Hwang, C., Cui, W., Xiong, Y., Yang, Z., Liu, Z., Hu, H., Wang, Z., Salas, R., Jose, J., Ram, P., et al · 2023
Later among the works it cites.
Scaling Vision-Language Models with Sparse Mixture of Experts
Shen, S., Yao, Z., Li, C., Darrell, T., Keutzer, K., and He, Y · 2023
Later among the works it cites.
A Hybrid Tensor-Expert-Data Parallelism Approach to Optimize Mixture-of-Experts Training
Singh, S., Ruwase, O., Awan, A. A., Rajbhandari, S., He, Y., and Bhatele, A · 2023
Later among the works it cites.
SlimPajama: A 627B Token Cleaned and Deduplicated Version of RedPajama
Soboleva, D., Al-Khateeb, F., Myers, R., Steeves, J. R., Hestness, J., and Dey, N · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning
Rajbhandari, S., Ruwase, O., Rasley, J., Smith, S., and He, Y · 2021
Cited alongside, same era.
Scaling Vision with Sparse Mixture of Experts
Riquelme, C., Puigcerver, J., Mustafa, B., Neumann, M., Jenatton, R., Pinto, A. S., Keysers, D., and Houlsby, N · 2021
Cited alongside, same era.
WinoGrande: An Adversarial Winograd Schema Challenge at Scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2021
Cited alongside, same era.
VinVL: Revisiting Visual Representations in Vision-Language Models
Zhang, P., Li, X., Hu, X., Yang, J., Zhang, L., Wang, L., Choi, Y., and Gao, J · 2021
Cited alongside, same era.
GLaM: Efficient Scaling of Language Models with Mixture-of-Experts
Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., Zoph, B., Fedus, L., Bosma, M. P., Zhou, Z., Wang, T., Wang, Y. E., Webster, K., Pellat, M., Robinson, K., Meier-Hellstern, K. S., Duke, T., Dixon, L., Zhang, K., Le, Q. V., Wu, Y., Chen, Z., and Cui, C · 2022
Cited alongside, same era.
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2022
Cited alongside, same era.
FasterMoE: Modeling and Optimizing Training of Large-Scale Dynamic Pre-Trained Models
He, J., Zhai, J., Antunes, T., Wang, H., Luo, F., Shi, S., and Li, Q · 2022
Cited alongside, same era.
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models
Wang, S., Wei, J., Sabne, A., Davis, A., Ilbeyi, B., Hechtman, B., Chen, D., Murthy, K. S., Maggioni, M., Zhang, Q., Kumar, S., Guo, T., Xu, Y., and Zhou, Z · 2023
Later among the works it cites.
MoLE: Mixture of LoRA Experts
Wu, X., Huang, S., and Wei, F · 2023
Later among the works it cites.
EdgeMoE: Fast On-Device Inference of MoE-based Large Language Models
Yi, R., Guo, L., Wei, S., Zhou, A., Wang, S., and Xu, M · 2023
Later among the works it cites.
MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism
Zhang, Z., Yang, D., Xia, Y., Ding, L., Tao, D., Zhou, X., and Cheng, D · 2023
Later among the works it cites.
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M · 2023
Later among the works it cites.
Chen, S., Jie, Z., and Ma, L · 2024
Closest in time.
DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
Dai, D., Deng, C., Zhao, C., Xu, R., Gao, H., Chen, D., Li, J., Zeng, W., Yu, X., Wu, Y., et al · 2024
Closest in time.
Introducing DBRX: A New State-of-the-Art Open LLM, March 2024
Databricks · 2024
Closest in time.
SiDA: Sparsity-Inspired Data-Aware Serving for Efficient and Scalable Large Mixture-of-Experts Models
Du, Z., Li, S., Wu, Y., Jiang, X., Sun, J., Zheng, Q., Wu, Y., Li, A., Li, H., and Chen, Y · 2024
Closest in time.
Higher Layers Need More LoRA Experts
Gao, C., Chen, K., Rao, J., Sun, B., Liu, R., Peng, D., Zhang, Y., Guo, X., Yang, J., and Subrahmanian, V · 2024
Closest in time.
Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference
Hwang, R., Wei, J., Cao, S., Hwang, C., Tang, X., Cao, T., and Yang, M · 2024
Closest in time.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al · 2024
Closest in time.
OLMoE: Open Mixture-of-Experts Language Models
Muennighoff, N., Soldaini, L., Groeneveld, D., Lo, K., Morrison, J., Min, S., Shi, W., Walsh, P., Tafjord, O., Lambert, N., et al · 2024
Closest in time.
Splitwise: Efficient Generative LLM Inference Using Phase Splitting
Patel, P., Choukse, E., Zhang, C., Shah, A., Goiri, Í., Maleki, S., and Bianchini, R · 2024
Closest in time.
Qwen1.5-MoE: Matching 7B Model Performance with 1/3 Activated Parameters”, February 2024
Qwen · 2024
Closest in time.
Loongserve: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism
Wu, B., Liu, S., Zhong, Y., Sun, P., Liu, X., and Jin, X · 2024
Closest in time.
OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models
Xue, F., Zheng, Z., Fu, Y., Ni, J., Zheng, Z., Zhou, W., and You, Y · 2024
Closest in time.
Jamba: A Hybrid Transformer-Mamba Language Model
Lieber, O., Lenz, B., Bata, H., Cohen, G., Osin, J., Dalmedigos, I., Safahi, E., Meirom, S., Belinkov, Y., Shalev-Shwartz, S., et al · 2025
Closest in time.