Fetching the paper…
Reading the bibliography…
The emergence of large-scale Mixture of Experts (MoE) models represents a significant advancement in artificial intelligence, offering enhanced model capacity and computational efficiency through conditional computation.
Loramoe: Alleviating world knowledge forgetting in large language models via moe-style plugin
Dou, S., Zhou, E., Liu, Y., Gao, S., Shen, W., Xiong, L., Zhou, Y., Wang, X., Xi, Z., Fan, X., et al · 1945
Earlier work this paper cites.
Adaptive mixture of local expert
Jacobs, R., Jordan, M., Nowlan, S., and Hinton, G · 1991
Earlier work this paper cites.
The meta-pi network: building distributed knowledge representations for robust multisource pattern recognition
Hampshire, J., and Waibel, A · 1992
Earlier work this paper cites.
Quantum computing
Steane, A · 1998
Earlier work this paper cites.
Mixtures of gaussian processes
Tresp, V · 2000
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Dolan, B., and Brockett, C · 2005
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Roemmele, M., Bejan, C. A., and Gordon, A. S · 2011
Earlier work this paper cites.
Learning factored representations in a deep mixture of experts
Eigen, D., Ranzato, M., and Sutskever, I · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A. Y., and Potts, C · 2013
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al · 2015
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Rajpurkar, P · 2016
Earlier work this paper cites.
A survey of neuromorphic computing and neural networks in hardware
Schuman, C. D., Potok, T. E., Patton, R. M., Birdwell, J. D., Dean, M. E., Rose, G. S., and Plank, J. S · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A · 2017
Earlier work this paper cites.
Looking beyond the surface: A challenge set for reading comprehension over multiple sentences
Khashabi, D., Chaturvedi, S., Roth, M., Upadhyay, S., and Roth, D · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A · 2018
Earlier work this paper cites.
Wikiqa: A challenge dataset for open-domain question answering
Yang, Y., Yih, W.-t., and Meek, C · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W. W., Salakhutdinov, R., and Manning, C. D · 2018
Earlier work this paper cites.
Beyond data and model parallelism for deep neural networks
Jia, Z., Zaharia, M., and Aiken, A · 2019
Earlier work this paper cites.
Fastertransformer, 2019
NVIDIA · 2019
Earlier work this paper cites.
Fairseq, 2019
Research, F. A · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z · 2020
Earlier work this paper cites.
A mixture of h - 1 heads is better than h heads
Peng, H., Schwartz, R., Li, D., and Smith, N. A · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y · 2020
Earlier work this paper cites.
Mlperf inference benchmark
Reddi, V. J., Cheng, C., Kanter, D., Mattson, P., Schmuelling, G., Wu, C.-J., Anderson, B., Breughe, M., Charlebois, M., Chou, W., et al · 2020
Earlier work this paper cites.
Deep mixture of experts via shallow embedding
Wang, X., Yu, F., Dunlap, L., Ma, Y.-A., Wang, R., Mirhoseini, A., Darrell, T., and Gonzalez, J. E · 2020
Earlier work this paper cites.
Efficient large scale language modeling with mixtures of experts
Artetxe, M., Bhosale, S., Goyal, N., Mihaylov, T., Ott, M., Shleifer, S., Lin, X. V., Du, J., Iyer, S., Pasunuru, R., et al · 2021
Earlier work this paper cites.
Fastmoe: A fast mixture-of-expert training system
He, J., Qiu, J., Zeng, A., Yang, Z., Zhai, J., and Tang, J · 2021
Earlier work this paper cites.
Base layers: Simplifying training of large, sparse models
Lewis, M., Bhosale, S., Dettmers, T., Goyal, N., and Zettlemoyer, L · 2021
Earlier work this paper cites.
Align before fuse: Vision and language representation learning with momentum distillation
Li, J., Selvaraju, R., Gotmare, A., Joty, S., Xiong, C., and Hoi, S. C. H · 2021
Earlier work this paper cites.
Efficient large-scale language model training on gpu clusters using megatron-lm
Narayanan, D., Shoeybi, M., Casper, J., LeGresley, P., Patwary, M., Korthikanti, V., Vainbrand, D., Kashinkunti, P., Bernauer, J., Catanzaro, B., et al · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Hash layers for large sparse models
Roller, S., Sukhbaatar, S., Szlam, A., and Weston, J · 2021
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2021
Earlier work this paper cites.
Vit-yolo: Transformer-based yolo for object detection
Zhang, Z., Lu, X., Cao, G., Yang, Y., Jiao, L., and Liu, F · 2021
Earlier work this paper cites.
Revisiting neural scaling laws in language and vision
Alabdulmohsin, I. M., Neyshabur, B., and Zhai, X · 2022
Earlier work this paper cites.
Ta-moe: Topology-aware large scale mixture-of-expert training
Chen, C., Li, M., Wu, Z., Yu, D., and Yang, C · 2022
Earlier work this paper cites.
Task-specific expert pruning for sparse mixture-of-experts
Chen, T., Huang, S., Xie, Y., Jiao, B., Jiang, D., Zhou, H., Li, J., and Wei, F · 2022
Earlier work this paper cites.
No language left behind: Scaling human-centered machine translation
Costa-jussà, M. R., Cross, J., Çelebi, O., Elbayad, M., Heafield, K., Heffernan, K., Kalbassi, E., Lam, J., Licht, D., Maillard, J., et al · 2022
Earlier work this paper cites.
Glam: Efficient scaling of language models with mixture-of-experts
Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., et al · 2022
Earlier work this paper cites.
M3vit: Mixture-of-experts vision transformer for efficient multi-task learning with model-accelerator co-design
Fan, Z., Sarkar, R., Jiang, Z., Chen, T., Zou, K., Cheng, Y., Hao, C., Wang, Z., et al · 2022
Earlier work this paper cites.
A review of sparse expert models in deep learning
Fedus, W., Dean, J., and Zoph, B · 2022
Earlier work this paper cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2022
Earlier work this paper cites.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Frantar, E., Ashkboos, S., Hoefler, T., and Alistarh, D · 2022
Earlier work this paper cites.
Parameter-efficient mixture-of-experts architecture for pre-trained language models
Gao, Z.-F., Liu, P., Zhao, W. X., Lu, Z.-Y., and Wen, J.-R · 2022
Earlier work this paper cites.
Fastermoe: modeling and optimizing training of large-scale dynamic pre-trained models
He, J., Zhai, J., Antunes, T., Wang, H., Luo, F., Shi, S., and Li, Q · 2022
Earlier work this paper cites.
Who says elephants can’t run: Bringing large scale moe models into cloud scale production
Kim, Y. J., Henry, R., Fahim, R., and Awadalla, H. H · 2022
Earlier work this paper cites.
Branch-train-merge: Embarrassingly parallel training of expert language models
Li, M., Gururangan, S., Dettmers, T., Lewis, M., Althoff, T., Smith, N. A., and Zettlemoyer, L · 2022
Earlier work this paper cites.
Gating dropout: Communication-efficient regularization for sparsely activated transformers
Liu, R., Kim, Y. J., Muzio, A., and Hassan, H · 2022
Earlier work this paper cites.
Bagualu: targeting brain scale pretrained models with over 37 million cores
Ma, Z., He, J., Qiu, J., Cao, H., Wang, Y., Sun, Z., Zheng, L., Wang, H., Tang, S., Zheng, T., Lin, J., Feng, G., Huang, Z., Gao, J., Zeng, A., Zhang, J., Zhong, R., Shi, T., Liu, S., Zheng, W., Tang, J., Yang, H., Liu, X., Zhai, J., and Chen, W · 2022
Earlier work this paper cites.
Hetumoe: An efficient trillion-scale mixture-of-expert distributed training system
Nie, X., Zhao, P., Miao, X., Zhao, T., and Cui, B · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Carbon-aware computing for datacenters
Radovanović, A., Koningstein, R., Schneider, I., Chen, B., Duarte, A., Roy, B., Xiao, D., Haridasan, M., Hung, P., Care, N., et al · 2022
Earlier work this paper cites.
Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation AI scale
Rajbhandari, S., Li, C., Yao, Z., Zhang, M., Aminabadi, R. Y., Awan, A. A., Rasley, J., and He, Y · 2022
Earlier work this paper cites.
Knowledge distillation for mixture of experts models in speech recognition
Salinas, F. C., Kumatani, K., Gmyr, R., Liu, L., and Shi, Y · 2022
Earlier work this paper cites.
Opportunities for neuromorphic computing algorithms and applications
Schuman, C. D., Kulkarni, S. R., Parsa, M., Mitchell, J. P., Date, P., and Kay, B · 2022
Earlier work this paper cites.
Se-moe: A scalable and efficient mixture-of-experts distributed training and inference system
Shen, L., Wu, Z., Gong, W., Hao, H., Bai, Y., Wu, H., Wu, X., Bian, J., Xiong, H., Yu, D., et al · 2022
Earlier work this paper cites.
Vaqf: Fully automatic software-hardware co-design framework for low-bit vision transformer
Sun, M., Ma, H., Kang, G., Jiang, Y., Chen, T., Ma, X., Wang, Z., and Wang, Y · 2022
Earlier work this paper cites.
Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework
Wang, P., Yang, A., Men, R., Lin, J., Bai, S., Li, Z., Ma, J., Zhou, C., Zhou, J., and Yang, H · 2022
Earlier work this paper cites.
One student knows all experts know: From sparse to dense
Xue, F., He, X., Ren, X., Lou, Y., and You, Y · 2022
Earlier work this paper cites.
Accelerating large-scale distributed neural network training with spmd parallelism
Zhang, S., Diao, L., Wu, C., Wang, S., and Lin, W · 2022
Earlier work this paper cites.
Mixture of attention heads: Selecting attention heads per token
Zhang, X., Shen, Y., Huang, Z., Zhou, J., Rong, W., and Xiong, Z · 2022
Earlier work this paper cites.
Alpa: Automating inter- and Intra-Operator parallelism for distributed deep learning
Zheng, L., Li, Z., Zhang, H., Zhuang, Y., Chen, Z., Huang, Y., Wang, Y., Xu, Y., Zhuo, D., Xing, E. P., Gonzalez, J. E., and Stoica, I · 2022
Earlier work this paper cites.
Transpim: A memory-based acceleration via software-hardware co-design for transformer
Zhou, M., Xu, W., Kang, J., and Rosing, T · 2022
Earlier work this paper cites.
Mixture-of-experts with expert choice routing
Zhou, Y., Lei, T., Liu, H., Du, N., Huang, Y., Zhao, V., Dai, A. M., Le, Q. V., and Laudon, J · 2022
Earlier work this paper cites.
Gkd: Generalized knowledge distillation for auto-regressive sequence models
Agarwal, R., Vieillard, N., Stanczyk, P., Ramos, S., Geist, M., and Bachem, O · 2023
Earlier work this paper cites.
The claude 3 model family: Opus, sonnet, haiku, 2023
Anthropic · 2023
Earlier work this paper cites.
Reproducible scaling laws for contrastive language-image learning
Cherti, M., Beaumont, R., Wightman, R., Wortsman, M., Ilharco, G., Gordon, C., Schuhmann, C., Schmidt, L., and Jitsev, J · 2023
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2023
Earlier work this paper cites.
Patch-level routing in mixture-of-experts is provably sample-efficient for convolutional neural networks
Chowdhury, M. N. R., Zhang, S., Wang, M., Liu, S., and Chen, P.-Y · 2023
Earlier work this paper cites.
Switchhead: Accelerating transformers with mixture-of-experts attention
Csordás, R., Piękos, P., Irie, K., and Schmidhuber, J · 2023
Earlier work this paper cites.
Optimizing dynamic neural networks with brainstorm
Cui, W., Han, Z., Ouyang, L., Wang, Y., Zheng, N., Ma, L., Yang, Y., Yang, F., Xue, J., Qiu, L., Zhou, L., Chen, Q., Tan, H., and Guo, M · 2023
Earlier work this paper cites.
Scaling vision transformers to 22 billion parameters
Dehghani, M., Djolonga, J., Mustafa, B., Padlewski, P., Heek, J., Gilmer, J., Steiner, A. P., Caron, M., Geirhos, R., Alabdulmohsin, I., Jenatton, R., Beyer, L., Tschannen, M., Arnab, A., Wang, X., Riquelme Ruiz, C., Minderer, M., Puigcerver, J., Evci, U., Kumar, M., Steenkiste, S. V., Elsayed, G. F., Mahendran, A., Yu, F., Oliver, A., Huot, F., Bastings, J., Collier, M., Gritsenko, A. A., Birodkar, V., Vasconcelos, C. N., Tay, Y., Mensink, T., Kolesnikov, A., Pavetic, F., Tran, D., Kipf, T., Lucic, M., Zhai, X., Keysers, D., Harmsen, J. J., and Houlsby, N · 2023
Cited alongside, same era.
Fast inference of mixture-of-experts language models with offloading
Eliseev, A., and Mazur, D · 2023
Cited alongside, same era.
Llmcarbon: Modeling the end-to-end carbon footprint of large language models
Faiz, A., Kaneda, S., Wang, R., Osi, R., Sharma, P., Chen, F., and Jiang, L · 2023
Cited alongside, same era.
Qmoe: Practical sub-1-bit compression of trillion-parameter models
Frantar, E., and Alistarh, D · 2023
Cited alongside, same era.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., de las Casas, D., Hanna, E. B., Bressand, F., Lengyel, G., Bour, G., Lample, G., Lavaud, L. R., Saulnier, L., Lachaux, M.-A., Stock, P., Subramanian, S., Yang, S., Antoniak, S., Scao, T. L., Gervet, T., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2024
Closest in time.
Moe++: Accelerating mixture-of-experts methods with zero-computation experts
Jin, P., Zhu, B., Yuan, L., and Yan, S · 2024
Closest in time.
Moh: Multi-head attention as mixture-of-head attention
Jin, P., Zhu, B., Yuan, L., and Yan, S · 2024
Closest in time.
Fiddler: Cpu-gpu orchestration for fast inference of mixture-of-experts models
Kamahori, K., Gu, Y., Zhu, K., and Kasikci, B · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Llama.cpp, 2023
Gerganov, G · 2023
Cited alongside, same era.
Knowledge distillation of large language models
Gu, Y., Dong, L., Wei, F., and Huang, M · 2023
Cited alongside, same era.
Merging experts into one: Improving computational efficiency of mixture of experts
He, S., Fan, R.-Z., Ding, L., Shen, L., Zhou, T., and Tao, D · 2023
Cited alongside, same era.
Towards moe deployment: Mitigating inefficiencies in mixture-of-expert (moe) inference
Huang, H., Ardalani, N., Sun, A., Ke, L., Lee, H.-H. S., Sridhar, A., Bhosale, S., Wu, C.-J., and Lee, B · 2023
Cited alongside, same era.
Experts weights averaging: A new general training scheme for vision transformers
Huang, Y., Ye, P., Huang, X., Li, S., Chen, T., He, T., and Ouyang, W · 2023
Cited alongside, same era.
Llm-perf leaderboard, 2023
Huggingface · 2023
Cited alongside, same era.
Transformers: State-of-the-art machine learning for jax, pytorch and tensorflow, 2023
Huggingface · 2023
Cited alongside, same era.
Tutel: Adaptive mixture-of-experts at scale
Hwang, C., Cui, W., Xiong, Y., Yang, Z., Liu, Z., Hu, H., Wang, Z., Salas, R., Jose, J., Ram, P., et al · 2023
Cited alongside, same era.
Kim, S., Kim, Y., Moon, K., and Jang, M · 2024
Closest in time.
Monde: Mixture of near-data experts for large-scale sparse models
Kim, T., Choi, K., Cho, Y., Cho, J., Lee, H.-J., and Sim, J · 2024
Closest in time.
Stun: Structured-then-unstructured pruning for scalable moe pruning
Lee, J., Qiao, A., Campos, D. F., Yao, Z., He, Y., et al · 2024
Closest in time.
Llm inference serving: Survey of recent advances and opportunities
Li, B., Jiang, Y., Gadepally, V., and Tiwari, D · 2024
Closest in time.
Locmoe: A low-overhead moe for large language model training
Li, J., Sun, Z., He, X., Zeng, L., Lin, Y., Li, E., Zheng, B., Zhao, R., and Chen, X · 2024
Closest in time.
Optimizing mixture-of-experts inference time combining model deployment and communication scheduling
Li, J., Tripathi, S., Rastogi, L., Lei, Y., Pan, R., and Xia, Y · 2024
Closest in time.
Large language model inference acceleration: A comprehensive hardware perspective
Li, J., Xu, J., Huang, S., Chen, Y., Li, W., Liu, J., Lian, Y., Pan, J., Ding, L., Zhou, H., et al · 2024
Closest in time.
Examining post-training quantization for mixture-of-experts: A benchmark
Li, P., Jin, X., Cheng, Y., and Chen, T · 2024
Closest in time.
Jamba: A hybrid transformer-mamba language model
Lieber, O., Lenz, B., Bata, H., Cohen, G., Osin, J., Dalmedigos, I., Safahi, E., Meirom, S., Belinkov, Y., Shalev-Shwartz, S., Abend, O., Alon, R., Asida, T., Bergman, A., Glozman, R., Gokhman, M., Manevich, A., Ratner, N., Rozen, N., Shwartz, E., Zusman, M., and Shoham, Y · 2024
Closest in time.
Flame: Fully leveraging moe sparsity for transformer on fpga
Lin, X., Tian, H., Xue, W., Ma, L., Cao, J., Zhang, M., Yu, J., and Wang, K · 2024
Closest in time.
Liu, E., Zhu, J., Lin, Z., Ning, X., Blaschko, M. B., Yan, S., Dai, G., Yang, H., and Wang, Y · 2024
Closest in time.
Liu, L., Kim, Y. J., Wang, S., Liang, C., Shen, Y., Cheng, H., Liu, X., Tanaka, M., Wu, X., Hu, W., et al · 2024
Closest in time.
Adamole: Fine-tuning large language models with adaptive mixture of low-rank adaptation experts, 2024
Liu, Z., and Luo, J · 2024
Closest in time.
Lu, X., Liu, Q., Xu, Y., Zhou, A., Huang, S., Zhang, B., Yan, J., and Li, H · 2024
Closest in time.
Luo, T., Lei, J., Lei, F., Liu, W., He, S., Zhao, J., and Liu, K · 2024
Closest in time.
Green edge ai: A contemporary survey
Mao, Y., Yu, X., Huang, K., Zhang, Y.-J. A., and Zhang, J · 2024
Closest in time.
Shortgpt: Layers in large language models are more redundant than you expect
Men, X., Xu, M., Zhang, Q., Wang, B., Lin, H., Lu, Y., Han, X., and Chen, W · 2024
Closest in time.
Olmoe: Open mixture-of-experts language models
Muennighoff, N., Soldaini, L., Groeneveld, D., Lo, K., Morrison, J., Min, S., Shi, W., Walsh, P., Tafjord, O., Lambert, N., Gu, Y., Arora, S., Bhagia, A., Schwenk, D., Wadden, D., Wettig, A., Hui, B., Dettmers, T., Kiela, D., Farhadi, A., Smith, N. A., Koh, P. W., Singh, A., and Hajishirzi, H · 2024
Closest in time.
Seer-moe: Sparse expert efficiency through regularization for mixture-of-experts
Muzio, A., Sun, A., and He, C · 2024
Closest in time.
Libmoe: A library for comprehensive benchmarking mixture of experts in large language models
Nguyen, N. V., Doan, T. T., Tran, L., Nguyen, V., and Pham, Q · 2024
Closest in time.
Photonic-electronic integrated circuits for high-performance computing and ai accelerators
Ning, S., Zhu, H., Feng, C., Gu, J., Jiang, Z., Ying, Z., Midkiff, J., Jain, S., Hlaing, M. H., Pan, D. Z., et al · 2024
Closest in time.
Quantum computing and ai in the cloud
Padmanaban, H · 2024
Closest in time.
Dense training, sparse inference: Rethinking training of mixture-of-experts language models
Pan, B., Shen, Y., Liu, H., Mishra, M., Zhang, G., Oliva, A., Raffel, C., and Panda, R · 2024
Closest in time.
Parm: Efficient training of large sparsely-activated models with dedicated schedules
Pan, X., Lin, W., Shi, S., Chu, X., Sun, W., and Li, B · 2024
Closest in time.
20.8 space-mate: A 303.5mw real-time sparse mixture-of-experts-based nerf-slam processor for mobile spatial computing
Park, G., Song, S., Sang, H., Im, D., Han, D., Kim, S., Lee, H., and Yoo, H.-J · 2024
Closest in time.
Learning more generalized experts by merging experts in mixture-of-experts
Park, S · 2024
Closest in time.
Eps-moe: Expert pipeline scheduler for cost-efficient moe inference
Qian, Y., Li, F., Ji, X., Zhao, X., Tan, J., Zhang, K., and Cai, X · 2024
Closest in time.
Green and sustainable ai research: an integrated thematic and topic modeling analysis
Raman, R., Pattnaik, D., Lathabai, H. H., Kumar, C., Govindan, K., and Nedungadi, P · 2024
Closest in time.
Enabling large dynamic neural network training with learning-based memory management
Ren, J., Xu, D., Yang, S., Zhao, J., Li, Z., Navasca, C., Wang, C., Xu, H., and Li, D · 2024
Closest in time.
Snowflake arctic: The best llm for enterprise ai — efficiently intelligent, truly open, April 2024
Research, S. A · 2024
Closest in time.
Revisiting smoe language models by evaluating inefficiencies with task specific expert pruning
Sarkar, S., Lausen, L., Cevher, V., Zha, S., Brox, T., and Karypis, G · 2024
Closest in time.
Jetmoe: Reaching llama2 performance with 0.1 m dollars
Shen, Y., Guo, Z., Cai, T., and Qin, Z · 2024
Closest in time.
Schemoe: An extensible mixture-of-experts distributed training system with tasks scheduling
Shi, S., Pan, X., Wang, Q., Liu, C., Ren, X., Hu, Z., Yang, Y., Li, B., and Chu, X · 2024
Closest in time.
Llava-mod: Making llava tiny via moe knowledge distillation
Shu, F., Liao, Y., Zhuo, L., Xu, C., Zhang, L., Zhang, G., Shi, H., Chen, L., Zhong, T., He, W., et al · 2024
Closest in time.
Mixture of cache-conditional experts for efficient mobile device inference
Skliar, A., van Rozendaal, T., Lepert, R., Boinovski, T., van Baalen, M., Nagel, M., Whatmough, P., and Bejnordi, B. E · 2024
Closest in time.
Promoe: Fast moe-based llm serving using proactive caching
Song, X., Zhong, Z., and Chen, R · 2024
Closest in time.
Branch-train-mix: Mixing expert llms into a mixture-of-experts llm
Sukhbaatar, S., Golovneva, O., Sharma, V., Xu, H., Lin, X. V., Rozière, B., Kahn, J., Li, D., Yih, W.-t., Weston, J., et al · 2024
Closest in time.
Hobbit: A mixed precision expert offloading system for fast moe inference
Tang, P., Liu, J., Hou, X., Pu, Y., Wang, J., Heng, P.-A., Li, C., and Guo, M · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Team, G., Georgiev, P., Lei, V. I., Burnell, R., Bai, L., Gulati, A., Tanzer, G., Vincent, D., Pan, Z., Wang, S., et al · 2024
Closest in time.
Tensorflow, 2024
Team, G. B · 2024
Closest in time.
Qwen1.5-moe: Matching 7b model performance with 1/3 activated parameters, February 2024
Team, Q · 2024
Closest in time.
Introducing dbrx: A new state-of-the-art open llm, March 2024
Team, T. M. R · 2024
Closest in time.
The evolution of mixture of experts: A survey from basics to breakthroughs
Vats, A., Raja, R., Jain, V., and Chadha, A · 2024
Closest in time.
Svd-llm: Truncation-aware singular value decomposition for large language model compression
Wang, X., Zheng, Y., Wan, Z., and Zhang, M · 2024
Closest in time.
Skywork-moe: A deep dive into training techniques for mixture-of-experts language models
Wei, T., Zhu, B., Zhao, L., Cheng, C., Li, B., Lü, W., Cheng, P., Zhang, J., Zhang, X., Zeng, L., Wang, X., Ma, Y., Hu, R., Yan, S., Fang, H., and Zhou, Y · 2024
Closest in time.
Yuan 2.0-m32: Mixture of experts with attention router
Wu, S., Luo, J., Chen, X., Li, L., Zhao, X., Yu, T., Wang, C., Wang, Y., Wang, F., Qiao, W., He, H., Zhang, Z., Sun, Z., Mao, J., and Shen, C · 2024
Closest in time.
Lazarus: Resilient and elastic training of mixture-of-experts models with adaptive expert placement
Wu, Y., Qu, W., Tao, T., Wang, Z., Bai, W., Li, Z., Tian, Y., Zhang, J., Lentz, M., and Zhuo, D · 2024
Closest in time.
Open release of grok-1, March 2024
xAI · 2024
Closest in time.
Moe-pruner: Pruning mixture-of-experts large language model using the hints from its router
Xie, Y., Zhang, Z., Zhou, D., Xie, C., Song, Z., Liu, X., Wang, Y., Lin, X., and Xu, A · 2024
Closest in time.
A survey of resource-efficient llm and multimodal foundation models
Xu, M., Yin, W., Cai, D., Yi, R., Xu, D., Wang, Q., Wu, B., Zhao, Y., Yang, C., Wang, S., et al · 2024
Closest in time.
Openmoe: An early effort on open mixture-of-experts language models
Xue, F., Zheng, Z., Fu, Y., Ni, J., Zheng, Z., Zhou, W., and You, Y · 2024
Closest in time.
Moe-infinity: Activation-aware expert offloading for efficient moe serving
Xue, L., Fu, Y., Lu, Z., Mai, L., and Marina, M · 2024
Closest in time.
Yang, A., Yang, B., Hui, B., Zheng, B., Yu, B., Zhou, C., Li, C., Li, C., Liu, D., Huang, F., Dong, G., Wei, H., Lin, H., Tang, J., Wang, J., Yang, J., Tu, J., Zhang, J., Ma, J., Yang, J., Xu, J., Zhou, J., Bai, J., He, J., Lin, J., Dang, K., Lu, K., Chen, K., Yang, K., Li, M., Xue, M., Ni, N., Zhang, P., Wang, P., Peng, R., Men, R., Gao, R., Lin, R., Wang, S., Bai, S., Tan, S., Zhu, T., Li, T., Liu, T., Ge, W., Deng, X., Zhou, X., Ren, X., Zhang, X., Wei, X., Ren, X., Liu, X., Fan, Y., Yao, Y., Zhang, Y., Wan, Y., Chu, Y., Liu, Y., Cui, Z., Zhang, Z., Guo, Z., and Fan, Z · 2024
Closest in time.
Yang, C., Sui, Y., Xiao, J., Huang, L., Gong, Y., Duan, Y., Jia, W., Yin, M., Cheng, Y., and Yuan, B · 2024
Closest in time.
Xmoe: Sparse models with fine-grained and adaptive expert selection
Yang, Y., Qi, S., Gu, W., Wang, C., Gao, C., and Xu, Z · 2024
Closest in time.
Exploiting inter-layer expert affinity for accelerating mixture-of-experts model inference
Yao, J., Anthony, Q., Shafi, A., Subramoni, H., and Panda, D. K. D · 2024
Closest in time.
Moesys: A distributed and efficient mixture-of-experts training and inference system for internet services
Yu, D., Shen, L., Hao, H., Gong, W., Wu, H., Bian, J., Dai, L., and Xiong, H · 2024
Closest in time.
Llm inference unveiled: Survey and roofline model insights
Yuan, Z., Shang, Y., Zhou, Y., Dong, Z., Zhou, Z., Xue, C., Wu, B., Li, Z., Gu, Q., Lee, Y. J., et al · 2024
Closest in time.
Duplex: A device for large language models with mixture of experts, grouped query attention, and continuous batching
Yun, S., Kyung, K., Cho, J., Choi, J., Kim, J., Kim, B., Lee, S., Sohn, K., and Ahn, J. H · 2024
Closest in time.
Bam! just like that: Simple and efficient parameter upcycling for mixture of experts
Zhang, Q., Gritsch, N., Gnaneshwar, D., Guo, S., Cairuz, D., Venkitesh, B., Foerster, J., Blunsom, P., Ruder, S., Ustun, A., et al · 2024
Closest in time.
Diversifying the expert knowledge for task-agnostic pruning in sparse mixture-of-experts
Zhang, Z., Liu, X., Cheng, H., Xu, C., and Gao, J · 2024
Closest in time.
Mpmoe: Memory efficient moe for pre-trained models with adaptive pipeline parallelism
Zhang, Z., Xia, Y., Wang, H., Yang, D., Hu, C., Zhou, X., and Cheng, D · 2024
Closest in time.
HyperMoE: Towards better mixture of experts via transferring among experts
Zhao, H., Qiu, Z., Wu, H., Wang, Z., He, Z., and Fu, J · 2024
Closest in time.
Adapmoe: Adaptive sensitivity-based expert gating and management for efficient moe inference
Zhong, S., Liang, L., Wang, Y., Wang, R., Huang, R., and Li, M · 2024
Closest in time.
Lima: Less is more for alignment
Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L., et al · 2024
Closest in time.
A survey on efficient inference for large language models
Zhou, Z., Ning, X., Hong, K., Fu, T., Xu, J., Li, S., Lou, Y., Wang, L., Yuan, Z., Li, X., et al · 2024
Closest in time.
Lightening-transformer: A dynamically-operated optically-interconnected photonic transformer accelerator
Zhu, H., Gu, J., Wang, H., Jiang, Z., Zhang, Z., Tang, R., Feng, C., Han, S., Chen, R. T., and Pan, D. Z · 2024
Closest in time.
Llama-moe: Building mixture-of-experts from llama with continual pre-training
Zhu, T., Qu, X., Dong, D., Ruan, J., Tong, J., He, C., and Cheng, Y · 2024
Closest in time.
Litemoe: Customizing on-device llm serving via proxy submodel tuning
Zhuang, Y., Zheng, Z., Wu, F., and Chen, G · 2024
Closest in time.
Optimizing distributed deployment of mixture-of-experts model inference in serverless computing
Liu, M., Wang, W., and Wu, C · 2025
Closest in time.
Minimax-01: Scaling foundation models with lightning attention, 2025
MiniMax, Li, A., Gong, B., Yang, B., Shan, B., Liu, C., Zhu, C., Zhang, C., Guo, C., Chen, D., Li, D., Jiao, E., Li, G., Zhang, G., Sun, H., Dong, H., Zhu, J., Zhuang, J., Song, J., Zhu, J., Han, J., Li, J., Xie, J., Xu, J., Yan, J., Zhang, K., Xiao, K., Kang, K., Han, L., Wang, L., Yu, L., Feng, L., Zheng, L., Chai, L., Xing, L., Ju, M., Chi, M., Zhang, M., Huang, P., Niu, P., Li, P., Zhao, P., Yang, Q., Xu, Q., Wang, Q., Wang, Q., Li, Q., Leng, R., Shi, S., Yu, S., Li, S., Zhu, S., Huang, T., Liang, T., Sun, W., Sun, W., Cheng, W., Li, W., Song, X., Su, X., Han, X., Zhang, X., Hou, X., Min, X., Zou, X., Shen, X., Gong, Y., Zhu, Y., Zhou, Y., Zhong, Y., Hu, Y., Fan, Y., Yu, Y., Yang, Y., Li, Y., Huang, Y., Li, Y., Huang, Y., Xu, Y., Mao, Y., Li, Z., Li, Z., Tao, Z., Ying, Z., Cong, Z., Qin, Z., Fan, Z., Yu, Z., Jiang, Z., and Wu, Z · 2025
Closest in time.
Efficient inference offloading for mixture-of-experts large language models in internet of medical things
Yuan, X., Kong, W., Luo, Z., and Xu, M · 2077
Closest in time.