Fetching the paper…
Reading the bibliography…
As the era of big data arrives, traditional artificial intelligence algorithms have difficulty processing the demands of massive and diverse data.
Estimation and feature selection in mixtures of generalized linear experts models
Huynh, B.T., Chamroukhi, F., 2019 · 1907
Earlier work this paper cites.
A logical calculus of the ideas immanent in nervous activity
McCulloch, W.S., Pitts, W., 1943 · 1943
Earlier work this paper cites.
A general black box theory
Bunge, M., 1963 · 1963
Earlier work this paper cites.
Artificial intelligence
Winston, P.H., 1984 · 1984
Earlier work this paper cites.
Induction of decision trees
Quinlan, J.R., 1986 · 1986
Earlier work this paper cites.
Adaptive mixtures of local experts
Jacobs, R.A., Jordan, M.I., Nowlan, S.J., Hinton, G.E., 1991 · 1991
Earlier work this paper cites.
Neural networks and their applications
Bishop, C.M., 1994 · 1994
Earlier work this paper cites.
Hierarchical mixtures of experts and the em algorithm
Jordan, M.I., Jacobs, R.A., 1994 · 1994
Earlier work this paper cites.
Reinforcement learning: A survey
Kaelbling, L.P., Littman, M.L., Moore, A.W., 1996 · 1996
Earlier work this paper cites.
Data preparation for data mining
Zhang, S., Zhang, C., Yang, Q., 2003 · 2003
Earlier work this paper cites.
Product of gaussians for speech recognition
Gales, M.J., Airey, S., 2006 · 2006
Earlier work this paper cites.
A survey on transfer learning
Pan, S.J., Yang, Q., 2009 · 2009
Earlier work this paper cites.
Federated learning using a mixture of experts
Zec, E.L., Mogren, O., Martinsson, J., Sütfeld, L.R., Gillblad, D., 2020 · 2010
Earlier work this paper cites.
Clustering high dimensional data
Assent, I., 2012 · 2012
Earlier work this paper cites.
Long short-term memory
Graves, A., Graves, A., 2012 · 2012
Earlier work this paper cites.
Advances in natural language processing
Hirschberg, J., Manning, C.D., 2015 · 2015
Earlier work this paper cites.
Deep learning
LeCun, Y., Bengio, Y., Hinton, G., 2015 · 2015
Earlier work this paper cites.
A critical review of recurrent neural networks for sequence learning
Lipton, Z.C., 2015 · 2015
Earlier work this paper cites.
A regularized root–quartic mixture of experts for complex classification problems
Abbasi, E., Shiri, M.E., Ghatee, M., 2016 · 2016
Earlier work this paper cites.
Learning to learn by gradient descent by gradient descent
Andrychowicz, M., Denil, M., Gomez, S., Hoffman, M.W., Pfau, D., Schaul, T., Shillingford, B., De Freitas, N., 2016 · 2016
Earlier work this paper cites.
Incremental learning algorithms and applications, in: European Symposium on Artificial Neural Networks
Gepperth, A., Hammer, B., 2016 · 2016
Earlier work this paper cites.
Edge computing: Vision and challenges
Shi, W., Cao, J., Zhang, Q., Li, Y., Xu, L., 2016 · 2016
Earlier work this paper cites.
Data mining in distributed environment: a survey
Gan, W., Lin, J.C.W., Chao, H.C., Zhan, J., 2017 · 2017
Earlier work this paper cites.
Dynamic routing between capsules
Sabour, S., Frosst, N., Hinton, G.E., 2017 · 2017
Earlier work this paper cites.
Outrageously Large Neural Networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., Dean, J., 2017 · 2017
Earlier work this paper cites.
Mixtures of regression models with incomplete and noisy data
Jung, B.C., Cheon, S., Lim, H.K., 2018 · 2018
Earlier work this paper cites.
Deep learning for computer vision: A brief review
Voulodimos, A., Doulamis, N., Doulamis, A., Protopapadakis, E., 2018 · 2018
Earlier work this paper cites.
Multi-source heterogeneous data fusion, in: 2018 International Conference on Artificial Intelligence and Big Data, IEEE. pp. 47–51
Zhang, L., Xie, Y., Xidao, L., Zhang, X., 2018 · 2018
Earlier work this paper cites.
A multimodal and hybrid deep neural network model for remaining useful life estimation
Al-Dulaimi, A., Zabihi, S., Asif, A., Mohammadi, A., 2019 · 2019
Earlier work this paper cites.
Regularized estimation and feature selection in mixtures of gaussian-gated experts models, in: Research School on Statistics and Data Science. Springer, pp. 42–56
Chamroukhi, F., Lecocq, F., Nguyen, H.D., 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding, in: The Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 4171–4186
Devlin, J., Chang, M., Lee, K., Toutanova, K., 2019 · 2019
Earlier work this paper cites.
Breaking the gridlock in mixture-of-experts: Consistent and efficient algorithms, in: International Conference on Machine Learning, PMLR. pp. 4304–4313
Makkuva, A., Viswanath, P., Kannan, S., Oh, S., 2019 · 2019
Earlier work this paper cites.
A mixture-of-experts model for vehicle prediction using an online learning approach, in: International Conference on Artificial Neural Networks, Springer. pp. 456–471
Mirus, F., Stewart, T.C., Eliasmith, C., Conradt, J., 2019 · 2019
Earlier work this paper cites.
Towards understanding knowledge distillation, in: International Conference on Machine Learning, PMLR. pp. 5142–5151
Phuong, M., Lampert, C., 2019 · 2019
Earlier work this paper cites.
Leveraging user comments for recommendation in e-commerce
Chu, P.M., Mao, Y.S., Lee, S.J., Hou, C.L., 2020 · 2020
Earlier work this paper cites.
Model compression and hardware acceleration for neural networks: A comprehensive survey
Deng, L., Li, G., Han, S., Shi, L., Xie, Y., 2020 · 2020
Earlier work this paper cites.
A survey on ensemble learning
Dong, X., Yu, Z., Cao, W., Shi, Y., Ma, Q., 2020 · 2020
Earlier work this paper cites.
Privacy-preserving distributed optimization via subspace perturbation: A general framework
Li, Q., Heusdens, R., Christensen, M.G., 2020 · 2020
Earlier work this paper cites.
Machine learning algorithms-a review
Mahesh, B., 2020 · 2020
Earlier work this paper cites.
MDLB: a metadata dynamic load balancing mechanism based on reinforcement learning
Wu, Z.q., Wei, J., Zhang, F., Guo, W., Xie, G.w., 2020 · 2020
Earlier work this paper cites.
Continuous action reinforcement learning from a mixture of interpretable experts
Akrour, R., Tateo, D., Peters, J., 2021 · 2021
Earlier work this paper cites.
Transfer learning based mixture of experts classification model for high-resolution remote sensing scene classification
Gong, X., Chen, Z., Wu, L., Xie, Z., Xu, Y., 2021 · 2021
Earlier work this paper cites.
DSelect-k: Differentiable selection in the mixture of experts with applications to multi-task learning
Hazimeh, H., Zhao, Z., Chowdhery, A., Sathiamoorthy, M., Chen, Y., Mazumder, R., Hong, L., Chi, E., 2021 · 2021
Earlier work this paper cites.
FastMoE: A fast mixture-of-expert training system
He, J., Qiu, J., Zeng, A., Yang, Z., Zhai, J., Tang, J., 2021 · 2021
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models, in: International Conference on Learning Representations
Hu, E.J., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al., 2021 · 2021
Earlier work this paper cites.
Review of multi-party secure computing research
Jiang, K., 2021 · 2021
Earlier work this paper cites.
Scalable and efficient MoE training for multitask multilingual models
Kim, Y.J., Awan, A.A., Muzio, A., Salinas, A.F.C., Lu, L., Hendy, A., Rajbhandari, S., He, Y., Awadalla, H.H., 2021 · 2021
Earlier work this paper cites.
Few-shot and continual learning with attentive independent mechanisms, in: IEEE/CVF international conference on computer vision, pp. 9455–9464
Lee, E., Huang, C.H., Lee, C.Y., 2021 · 2021
Earlier work this paper cites.
GShard: Scaling giant models with conditional computation and automatic sharding, in: International Conference on Learning Representations
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., Chen, Z., 2021 · 2021
Earlier work this paper cites.
A survey of convolutional neural networks: analysis, applications, and prospects
Li, Z., Liu, F., Yang, W., Peng, S., Zhou, J., 2021 · 2021
Earlier work this paper cites.
scMM: Mixture-of-experts multimodal deep generative model for single-cell multiomics data analysis
Minoura, K., Abe, K., Nam, H., Nishikawa, H., Shimamura, T., 2021 · 2021
Earlier work this paper cites.
A review on the attention mechanism of deep learning
Niu, Z., Zhong, G., Yu, H., 2021 · 2021
Earlier work this paper cites.
Scaling vision with sparse mixture of experts
Riquelme, C., Puigcerver, J., Mustafa, B., Neumann, M., Jenatton, R., Susano Pinto, A., Keysers, D., Houlsby, N., 2021 · 2021
Earlier work this paper cites.
KDExplainer: A task-oriented attention model for explaining knowledge distillation, in: The Thirtieth International Joint Conference on Artificial Intelligence, pp. 3228–3234
Xue, M., Song, J., Wang, X., Chen, Y., Wang, X., Song, M., 2021 · 2021
Earlier work this paper cites.
Lifelong mixture of variational autoencoders
Ye, F., Bors, A.G., 2021 · 2021
Cited alongside, same era.
Specializing versatile skill libraries using local mixture of experts, in: Conference on Robot Learning, PMLR. pp. 1423–1433
Celik, O., Zhou, D., Li, G., Becker, P., Neumann, G., 2022 · 2022
Cited alongside, same era.
Towards overcoming data scarcity in materials science: unifying models and datasets with a mixture of experts framework
Chang, R., Wang, Y.X., Ertekin, E., 2022 · 2022
Cited alongside, same era.
No Language Left Behind: Scaling human-centered machine translation
Costa-jussà, M.R., Cross, J., Çelebi, O., Elbayad, M., Heafield, K., Heffernan, K., Kalbassi, E., Lam, J., Licht, D., Maillard, J., et al., 2022 · 2022
Cited alongside, same era.
A sparse mixture-of-experts model with screening of genetic associations to guide disease subtyping
Courbariaux, M., De Santiago, K., Dalmasso, C., Danjou, F., Bekadar, S., Corvol, J.C., Martinez, M., Szafranski, M., Ambroise, C., 2022 · 2022
Zhou, J., Chen, Z., Wang, B., Huang, M., 2023 · 2023
Later among the works it cites.
Detrs with collaborative hybrid assignments training, in: The IEEE/CVF International Conference on Computer Vision, pp. 6748–6758
Zong, Z., Song, G., Liu, Y., 2023 · 2023
Later among the works it cites.
Mixture of tokens: Continuous moe through cross-example aggregation, in: The Thirty-eighth Annual Conference on Neural Information Processing Systems, pp. 1–24
Antoniak, S., Krutul, M., Pióro, M., Krajewski, J., Ludziejewski, J., Ciebiera, K., Król, K., Odrzygóźdź, T., Cygan, M., Jaszczur, S., 2024 · 2024
Later among the works it cites.
Arabpour, R., Armstrong, J., Galimberti, L., Kratsios, A., Livieri, G., 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
GLaM: Efficient scaling of language models with mixture-of-experts, in: International Conference on Machine Learning, PMLR. pp. 5547–5569
Du, N., Huang, Y., Dai, A.M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A.W., Firat, O., et al., 2022 · 2022
Cited alongside, same era.
M 3 ViT: Mixture-of-experts vision transformer for efficient multi-task learning with model-accelerator co-design
Fan, Z., Sarkar, R., Jiang, Z., Chen, T., Zou, K., Cheng, Y., Hao, C., Wang, Z., et al., 2022 · 2022
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., Shazeer, N., 2022 · 2022
Cited alongside, same era.
Sparsely activated mixture-of-experts are robust multi-task learners
Gupta, S., Mukherjee, S., Subudhi, K., Gonzalez, E., Jose, D., Awadallah, A.H., Gao, J., 2022 · 2022
Cited alongside, same era.
Text2Human: Text-driven controllable human image generation
Jiang, Y., Yang, S., Qiu, H., Wu, W., Loy, C.C., Liu, Z., 2022 · 2022
Cited alongside, same era.
Learning to adapt clinical sequences with residual mixture of experts, in: 20th International Conference on Artificial Intelligence in Medicine, pp. 155–166
Lee, J.M., Hauskrecht, M., 2022 · 2022
Cited alongside, same era.
Swin transformer v2: Scaling up capacity and resolution, in: The IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12009–12019
Liu, Z., Hu, H., Lin, Y., Yao, Z., Xie, Z., Wei, Y., Ning, J., Cao, Y., Zhang, Z., Dong, L., et al., 2022 · 2022
Cited alongside, same era.
Hybrid session-aware recommendation with feature-based models
Bauer, J., Jannach, D., 2024 · 2024
Later among the works it cites.
SUTRA: Scalable multilingual language model architecture
Bendale, A., Sapienza, M., Ripplinger, S., Gibbs, S., Lee, J., Mistry, P., 2024 · 2024
Later among the works it cites.
DeepSeekMoE: Towards ultimate expert specialization in mixture-of-experts language models, in: The 62nd Annual Meeting of the Association for Computational Linguistics, pp. 1280–1297
Dai, D., Deng, C., Zhao, C., Xu, R., Gao, H., Chen, D., Li, J., Zeng, W., Yu, X., Wu, Y., et al., 2024 · 2024
Later among the works it cites.
Redesigning multi-scale neural network for crowd counting
Du, Z., Shi, M., Deng, J., Zafeiriou, S., 2024 · 2024
Later among the works it cites.
FairMOE: counterfactually-fair mixture of experts with levels of interpretability
Germino, J., Moniz, N., Chawla, N.V., 2024 · 2024
Later among the works it cites.
Dynamic mixture of experts: An auto-tuning approach for efficient transformer models
Guo, Y., Cheng, Z., Tang, X., Tu, Z., Lin, T., 2024 · 2024
Later among the works it cites.
Offline reinforcement learning for mixture-of-expert dialogue management
Gupta, D., Chow, Y., Tulepbergenov, A., Ghavamzadeh, M., Boutilier, C., 2024 · 2024
Later among the works it cites.
DBRX: Creating an llm from scratch using databricks, in: Databricks Data Intelligence Platform: Unlocking the GenAI Revolution. Springer, pp. 311–330
Gupta, N., Yip, J., 2024 · 2024
Later among the works it cites.
Contrastive learning and mixture of experts enables precise vector embeddings
Hallee, L., Kapur, R., Patel, A., Gleghorn, J.P., Khomtchouk, B., 2024 · 2024
Later among the works it cites.
FuseMoE: Mixture-of-experts transformers for fleximodal fusion
Han, X., Nguyen, H., Harris, C., Ho, N., Saria, S., 2024 · 2024
Later among the works it cites.
He, X.O., 2024 · 2024
Later among the works it cites.
Pre-gated MoE: An algorithm-system co-design for fast and scalable mixture-of-expert inference, in: ACM/IEEE 51st Annual International Symposium on Computer Architecture, IEEE. pp. 1018–1031
Hwang, R., Wei, J., Cao, S., Hwang, C., Tang, X., Cao, T., Yang, M., 2024 · 2024
Later among the works it cites.
Mixture of experts with mixture of precisions for tuning quality of service
Imani, H., Amirany, A., El-Ghazawi, T., 2024 · 2024
Later among the works it cites.
M4oE: A foundation model for medical multimodal image segmentation with mixture of experts, in: International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer. pp. 621–631
Jiang, Y., Shen, Y., 2024 · 2024
Later among the works it cites.
MoE++: Accelerating mixture-of-experts methods with zero-computation experts
Jin, P., Zhu, B., Yuan, L., Yan, S., 2024 · 2024
Later among the works it cites.
Fiddler: CPU-GPU orchestration for fast inference of mixture-of-experts models
Kamahori, K., Gu, Y., Zhu, K., Kasikci, B., 2024 · 2024
Later among the works it cites.
Customizing language models with instance-wise LoRA for sequential recommendation
Kong, X., Wu, J., Zhang, A., Sheng, L., Lin, H., Wang, X., He, X., 2024 · 2024
Later among the works it cites.
Large language models in law: A survey
Lai, J., Gan, W., Wu, J., Qi, Z., Yu, P.S., 2024 · 2024
Later among the works it cites.
Continual traffic forecasting via mixture of experts
Lee, S., Park, C., 2024 · 2024
Later among the works it cites.
LocMoE: A low-overhead moe for large language model training, in: Thirty-Third International Joint Conference on Artificial Intelligence, pp. 6377–6387
Li, J., Sun, Z., He, X., Zeng, L., Lin, Y., Li, E., Zheng, B., Zhao, R., Chen, X., 2024 · 2024
Later among the works it cites.
AdaMoLE: Fine-tuning large language models with adaptive mixture of low-rank adaptation experts
Liu, Z., Luo, J., 2024 · 2024
Later among the works it cites.
Luo, L., Wu, M., Li, M., Xin, Y., Wang, Q., Vardhanabhuti, V., Chu, W.C., Li, Z., Zhou, J., Rajpurkar, P., et al., 2024 · 2024
Later among the works it cites.
Parm: Efficient training of large sparsely-activated models with dedicated schedules, in: IEEE Conference on Computer Communications, IEEE. pp. 1880–1889
Pan, X., Lin, W., Shi, S., Chu, X., Sun, W., Li, B., 2024 · 2024
Later among the works it cites.
Learning more generalized experts by merging experts in mixture-of-experts
Park, S., 2024 · 2024
Later among the works it cites.
CompeteSMoE–effective training of sparse mixture of experts via competition
Pham, Q., Do, G., Nguyen, H., Nguyen, T., Liu, C., Sartipi, M., Nguyen, B.T., Ramasamy, S., Li, X., Hoi, S., et al., 2024 · 2024
Later among the works it cites.
DMoERM: Recipes of mixture-of-experts for effective reward modeling, in: Findings of the Association for Computational Linguistics, pp. 7006–7028
Quan, S., 2024 · 2024
Later among the works it cites.
Divide and not forget: Ensemble of selectively trained experts in continual learning
Rypeść, G., Cygert, S., Khan, V., Trzciński, T., Zieliński, B., Twardowski, B., 2024 · 2024
Later among the works it cites.
Block selective reprogramming for on-device training of vision transformers
Sarkar, S., Kundu, S., Zheng, K., Beerel, P.A., 2024 · 2024
Later among the works it cites.
Adaptive utilization of cross-scenario information for multi-scenario recommendation
Shu, X., Han, R., Li, X., Lin, W., 2024 · 2024
Later among the works it cites.
Mixture of cache-conditional experts for efficient mobile device inference
Skliar, A., van Rozendaal, T., Lepert, R., Boinovski, T., van Baalen, M., Nagel, M., Whatmough, P., Bejnordi, B.E., 2024 · 2024
Later among the works it cites.
Hunyuan-large: An open-source moe model with 52 billion activated parameters by tencent
Sun, X., Chen, Y., Huang, Y., Xie, R., Zhu, J., Zhang, K., Li, S., Yang, Z., Han, J., Shu, X., et al., 2024 · 2024
Later among the works it cites.
Web3: The next internet revolution
Wan, S., Lin, H., Gan, W., Chen, J., Yu, P.S., 2024 · 2024
Later among the works it cites.
Skywork-MoE: A deep dive into training techniques for mixture-of-experts language models
Wei, T., Zhu, B., Zhao, L., Cheng, C., Li, B., Lü, W., Cheng, P., Zhang, J., Zhang, X., Zeng, L., et al., 2024 · 2024
Later among the works it cites.
Addressing environmental stochasticity in reconfigurable intelligent surface aided unmanned aerial vehicle networks: Multi-task deep reinforcement learning based optimization for physical layer security
Wong, Y.J., Tham, M.L., Kwan, B.H., Iqbal, A., 2024 · 2024
Later among the works it cites.
Enhancing diversity for logical table-to-text generation with mixture of experts
Wu, J., Hou, M., 2024 · 2024
Later among the works it cites.
AutoEIS: Automatic feature embedding, interaction and selection on default prediction
Xiao, K., Jiang, X., Hou, P., Zhu, H., 2024 · 2024
Later among the works it cites.
MoE-Pruner: Pruning mixture-of-experts large language model using the hints from its router
Xie, Y., Zhang, Z., Zhou, D., Xie, C., Song, Z., Liu, X., Wang, Y., Lin, X., Xu, A., 2024 · 2024
Later among the works it cites.
Using mixture of experts to accelerate dataset distillation
Xu, Z., Fu, Z., 2024 · 2024
Later among the works it cites.
RAPHAEL: Text-to-image generation via large mixture of diffusion paths
Xue, Z., Song, G., Guo, Q., Liu, B., Zong, Z., Liu, Y., Luo, P., 2024 · 2024
Later among the works it cites.
Variational mixture of stochastic experts auto-encoder for multi-modal recommendation
Yi, J., Chen, Z., 2024 · 2024
Later among the works it cites.
Predicting energy consumption for hybrid energy systems toward sustainable manufacturing: A physics-informed approach using Pi-MMoE
Yuan, M., Liu, J., Chen, Z., Guo, Q., Yuan, M., Li, J., Yu, G., 2024 · 2024
Later among the works it cites.
Definition and analysis of gray matter atrophy subtypes in mild cognitive impairment based on data-driven methods
Zhang, B., Xu, M., Wu, Q., Ye, S., Zhang, Y., Li, Z., Initiative, A.D.N., 2024 · 2024
Later among the works it cites.
Large language models for medicine: a survey
Zheng, Y., Gan, W., Chen, Z., Qi, Z., Liang, Q., Yu, P.S., 2024 · 2024
Later among the works it cites.
Lory: Fully differentiable mixture-of-experts for autoregressive language model pre-training
Zhong, Z., Xia, M., Chen, D., Lewis, M., 2024 · 2024
Later among the works it cites.
An adaptive multi-scale feature fusion and adaptive mixture-of-experts multi-task model for industrial equipment health status assessment and remaining useful life prediction
Zhou, L., Wang, H., 2024 · 2024
Later among the works it cites.
Data scarcity in recommendation systems: A survey
Chen, Z., Gan, W., Wu, J., Hu, K., Lin, H., 2025 · 2025
Closest in time.
Artificial intelligence in landscape architecture: A survey
Xing, Y., Gan, W., Chen, Q., 2025 · 2025
Closest in time.
AUFormer: Vision transformers are parameter-efficient facial action unit detectors, in: European Conference on Computer Vision, Springer. pp. 427–445
Yuan, K., Yu, Z., Liu, X., Xie, W., Yue, H., Yang, J., 2025 · 2025
Closest in time.
Pacer and Runner: Cooperative learning framework between single-and cross-domain sequential recommendation, in: The 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 2071–2080
Park, C., Kim, T., Yoon, H., Hong, J., Yu, Y., Cho, M., Choi, M., Choo, J., 2024 · 2080
Closest in time.