Fetching the paper…
Reading the bibliography…
Mixture-of-experts (MoE) model incorporates the power of multiple submodels via gating functions to achieve greater performance in numerous regression and classification applications.
Maximum Likelihood from Incomplete Data Via the EM Algorithm
Dempster, A. P., Laird, N. M., and Rubin, D. B · 1977
Earlier work this paper cites.
Adaptive mixtures of local experts
Jacobs, R. A., Jordan, M. I., Nowlan, S. J., and Hinton, G. E · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the EM algorithm
Jordan, M. I. and Jacobs, R. A · 1994
Earlier work this paper cites.
Bayesian inference in mixtures-of-experts and hierarchical mixtures-of-experts models with an application to speech recognition
Peng, F., Jacobs, R., and Tanner, M · 1996
Earlier work this paper cites.
Assouad, Fano, and Le Cam
Yu, B · 1997
Earlier work this paper cites.
Improved learning algorithms for mixture of experts in multiclass classification
Chen, K., Xu, L., and Chi, H · 1999
Earlier work this paper cites.
Empirical Processes in M-estimation
van de Geer, S · 2000
Earlier work this paper cites.
Topics in Optimal Transportation
Villani, C · 2003
Earlier work this paper cites.
Identifiability of finite mixtures of multinomial logit models with varying and fixed effects
Grün, B. and Leisch, F · 2008
Earlier work this paper cites.
Optimal transport: Old and New
Villani, C · 2008
Earlier work this paper cites.
Variational mixture of experts for classification with applications to landmine detection
Yuksel, S. and Gader, P · 2010
Earlier work this paper cites.
High-dimensional regression with Gaussian mixtures and partially-latent response variables
Deleforge, A., Forbes, F., and Horaud, R · 2015
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J · 2017
Earlier work this paper cites.
Estimation and feature selection in mixtures of generalized linear experts models
Huynh, B. and Chamroukhi, F · 2019
Cited alongside, same era.
Global convergence of the em algorithm for mixtures of two component linear regression
Kwon, J., Qian, W., Caramanis, C., Chen, Y., and Davis, D · 2019
Cited alongside, same era.
Conformer: Convolution-augmented transformer for speech recognition
Gulati, A., Qin, J., Chiu, C., Parmar, N., Zhang, Y., Yu, J., Han, W., Wang, S., Zhang, Z., Wu, Y., and Pang, R · 2020
Cited alongside, same era.
EM converges for a mixture of many linear regressions
Kwon, J. and Caramanis, C · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Cited alongside, same era.
Sparsely activated mixture-of-experts are robust multi-task learners
Gupta, S., Mukherjee, S., Subudhi, K., Gonzalez, E., Jose, D., Awadallah, A., and Gao, J · 2022
Later among the works it cites.
Convergence rates for Gaussian mixtures of experts
Ho, N., Yang, C. Y., and Jordan, M. I · 2022
Later among the works it cites.
M 3 ViT: Mixture-of-experts vision transformer for efficient multi-task learning with model-accelerator co-design
Liang, H., Fan, Z., Sarkar, R., Jiang, Z., Chen, T., Zou, K., Cheng, Y., Hao, C., and Wang, Z · 2022
Later among the works it cites.
Refined convergence rates for maximum likelihood estimation under finite mixture models
Manole, T. and Ho, N · 2022
Later among the works it cites.
Evomoe: An evolutional mixture-of-experts training framework via dense-to-sparse gate
Nie, X., Miao, X., Cao, S., Ma, L., Liu, Q., Xue, J., Miao, Y., Liu, Y., Yang, Z., and Cui, B · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dselect-k: Differentiable selection in the mixture of experts with applications to multi-task learning
Hazimeh, H., Zhao, Z., Chowdhery, A., Sathiamoorthy, M., Chen, Y., Mazumder, R., Hong, L., and Chi, E. H · 2021
Cited alongside, same era.
On the minimax optimality of the em algorithm for learning two-component mixed linear regression
Kwon, J., Ho, N., and Caramanis, C · 2021
Cited alongside, same era.
GShard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z · 2021
Cited alongside, same era.
Probabilistic mixture-of-experts for efficient deep reinforcement learning
Ren, J., Li, Y., Ding, Z., Pan, W., and Dong, H · 2021
Cited alongside, same era.
Scaling vision with sparse mixture of experts
Riquelme, C., Puigcerver, J., Mustafa, B., Neumann, M., Jenatton, R., Pint, A. S., Keysers, D., and Houlsby, N · 2021
Cited alongside, same era.
VLMo: Unified vision-language pre-training with mixture-of-modality-experts
Bao, H., Wang, W., Dong, L., Liu, Q., Mohammed, O.-K., Aggarwal, K., Som, S., Piao, S., and Wei, F · 2022
Cited alongside, same era.
Glam: Efficient scaling of language models with mixture-of-experts
Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A., Firat, O., Zoph, B., Fedus, L., Bosma, M., Zhou, Z., Wang, T., Wang, E., Webster, K., Pellat, M., Robinson, K., Meier-Hellstern, K., Duke, T., Dixon, L., Zhang, K., Le, Q., Wu, Y., Chen, Z., and Cui, C · 2022
Cited alongside, same era.
Functional mixture-of-experts for classification
Pham, N. and Chamroukhi, F · 2022
Later among the works it cites.
Speechmoe2: Mixture-of-experts model with improved routing
You, Z., Feng, S., Su, D., and Yu, D · 2022
Later among the works it cites.
Mixture-of-experts with expert choice routing
Zhou, Y., Lei, T., Liu, H., Du, N., Huang, Y., Zhao, V., Dai, A. M., Chen, Z., Le, Q., and Laudon, J · 2022
Later among the works it cites.
Mod-squad: Designing mixtures of experts as modular multi-task learners
Chen, Z., Wang, P., Ma, L., Wong, K. K., and Wu, Q · 2023
Closest in time.
A mixture-of-expert approach to RL-based dialogue management
Chow, Y., Tulepbergenov, A., Nachum, O., Gupta, D., Ryu, M., Ghavamzadeh, M., and Boutilier, C · 2023
Closest in time.
HyperRouter: Towards efficient training and inference of sparse mixture of experts
Do, T., Le, H., Nguyen, T., Pham, Q., Nguyen, B., Doan, T., Liu, C., Ramasamy, S., Li, X., and HOI, S · 2023
Closest in time.
Demystifying softmax gating function in Gaussian mixture of experts
Nguyen, H., Nguyen, T., and Ho, N · 2023
Closest in time.
Brainformers: Trading simplicity for efficiency
Zhou, Y., Du, N., Huang, Y., Peng, D., Lan, C., Huang, D., Shakeri, S., So, D., Dai, A., Lu, Y., Chen, Z., Le, Q., Cui, C., Laudon, J., and Dean, J · 2023
Closest in time.