Fetching the paper…
Reading the bibliography…
We conduct the convergence analysis of parameter estimation in the contaminated mixture of experts.
Adaptive mixtures of local experts
Jacobs, R. A., Jordan, M. I., Nowlan, S. J., and Hinton, G. E. (1991) · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the EM algorithm
Jordan, M. I. and Jacobs, R. A. (1994) · 1994
Earlier work this paper cites.
Empirical Processes in M-estimation
van de Geer, S. (2000) · 2000
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J. (2017) · 2017
Earlier work this paper cites.
Parameter recovery in two-component contamination mixtures: The lˆ2 strategy
Gadat, S., Kahn, J., Marteau, C., and Maugis-Rabusseau, C. (2020) · 2020
Earlier work this paper cites.
DSelect-k: Differentiable Selection in the Mixture of Experts with Applications to Multi-Task Learning
Hazimeh, H., Zhao, Z., Chowdhery, A., Sathiamoorthy, M., Chen, Y., Mazumder, R., Hong, L., and Chi, E. (2021) · 2021
Earlier work this paper cites.
Scaling vision with sparse mixture of experts
Riquelme, C., Puigcerver, J., Mustafa, B., Neumann, M., Jenatton, R., Pint, A. S., Keysers, D., and Houlsby, N. (2021) · 2021
Earlier work this paper cites.
Scaling vision with sparse mixture of experts
Ruiz, C., Puigcerver, J., Mustafa, B., Neumann, M., Jenatton, R., Pinto, A., Keysers, D., and Houlsby, N. (2021) · 2021
Earlier work this paper cites.
Strong identifiability and parameter learning in regression with heterogeneous response
Do, D., Do, L., and Nguyen, X. (2022) · 2022
Earlier work this paper cites.
Glam: Efficient scaling of language models with mixture-of-experts
Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A., Firat, O., Zoph, B., Fedus, L., Bosma, M., Zhou, Z., Wang, T., Wang, E., Webster, K., Pellat, M., Robinson, K., Meier-Hellstern, K., Duke, T., Dixon, L., Zhang, K., Le, Q., Wu, Y., Chen, Z., and Cui, C. (2022) · 2022
Earlier work this paper cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N. (2022) · 2022
Cited alongside, same era.
Sparsely activated mixture-of-experts are robust multi-task learners
Gupta, S., Mukherjee, S., Subudhi, K., Gonzalez, E., Jose, D., Awadallah, A., and Gao, J. (2022) · 2022
Cited alongside, same era.
Convergence rates for Gaussian mixtures of experts
Ho, N., Yang, C.-Y., and Jordan, M. I. (2022) · 2022
Cited alongside, same era.
M 3 ViT: Mixture-of-Experts Vision Transformer for Efficient Multi-task Learning with Model-Accelerator Co-design
Liang, H., Fan, Z., Sarkar, R., Jiang, Z., Chen, T., Zou, K., Cheng, Y., Hao, C., and Wang, Z. (2022) · 2022
Cited alongside, same era.
A Mixture-of-Expert Approach to RL-based Dialogue Management
Chow, Y., Tulepbergenov, A., Nachum, O., Gupta, D., Ryu, M., Ghavamzadeh, M., and Boutilier, C. (2023) · 2023
Cited alongside, same era.
Brainformers: Trading simplicity for efficiency
Zhou, Y., Du, N., Huang, Y., Peng, D., Lan, C., Huang, D., Shakeri, S., So, D., Dai, A., Lu, Y., Chen, Z., Le, Q., Cui, C., Laudon, J., and Dean, J. (2023) · 2023
Later among the works it cites.
FuseMoE: Mixture-of-experts transformers for fleximodal fusion
Han, X., Nguyen, H., Harris, C., Ho, N., and Saria, S. (2024) · 2024
Closest in time.
Mixtral of experts
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., de las Casas, D., Hanna, E. B., Bressand, F., Lengyel, G., Bour, G., Lample, G., Lavaud, L. R., Saulnier, L., Lachaux, M.-A., Stock, P., Subramanian, S., Yang, S., Antoniak, S., Scao, T. L., Gervet, T., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E. (2024) · 2024
Closest in time.
Mixture of experts meets prompt-based continual learning
Le, M., Nguyen, A., Nguyen, H., Nguyen, T., Pham, T., Van Ngo, L., and Ho, N. (2024) · 2024
Closest in time.
Mixtures of experts unlock parameter scaling for deep RL
Obando Ceron, J. S., Sokar, G., Willi, T., Lyle, C., Farebrother, J., Foerster, J. N., Dziugaite, G. K., Precup, D., and Castro, P. S. (2024) · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Approximating two-layer feedforward networks for efficient transformers
Csordás, R., Irie, K., and Schmidhuber, J. (2023) · 2023
Cited alongside, same era.
Sparse mixture-of-experts are domain generalizable learners
Li, B., Shen, Y., Yang, J., Wang, Y., Ren, J., Che, T., Zhang, J., and Liu, Z. (2023) · 2023
Cited alongside, same era.
Botbuster: Multi-platform bot detection using a mixture of experts
Ng, L. H. X. and Carley, K. M. (2023) · 2023
Cited alongside, same era.
Demystifying softmax gating function in Gaussian mixture of experts
Nguyen, H., Nguyen, T., and Ho, N. (2023) · 2023
Cited alongside, same era.
Safe real-world autonomous driving by learning to predict and plan with a mixture of experts
Pini, S., Perone, C. S., Ahuja, A., Ferreira, A. S. R., Niendorf, M., and Zagoruyko, S. (2023) · 2023
Cited alongside, same era.
Convergence rates of parameter estimation for some weakly identifiable finite mixtures
Ho, N. and Nguyen, X. (2016a)
Cited in the paper.
On strong identifiability and convergence rates of parameter estimation in finite mixtures
Ho, N. and Nguyen, X. (2016b)
Cited in the paper.
Pham, Q., Do, G., Nguyen, H., Nguyen, T., Liu, C., Sartipi, M., Nguyen, B. T., Ramasamy, S., Li, X., Hoi, S., and Ho, N. (2024) · 2024
Closest in time.
From sparse to soft mixtures of experts
Puigcerver, J., Riquelme, C., Mustafa, B., and Houlsby, N. (2024) · 2024
Closest in time.
Mixture-of-experts meets instruction tuning: A winning combination for large language models
Shen, S., Hou, L., Zhou, Y., Du, N., Longpre, S., Wei, J., Chung, H. W., Zoph, B., Fedus, W., Chen, X., Vu, T., Wu, Y., Chen, W., Webson, A., Li, Y., Zhao, V. Y., Yu, H., Keutzer, K., Darrell, T., and Zhou, D. (2024) · 2024
Closest in time.
Pushing mixture of experts to the limit: Extremely parameter efficient moe for instruction tuning
Zadouri, T., Üstün, A., Ahmadian, A., Ermis, B., Locatelli, A., and Hooker, S. (2024) · 2024
Closest in time.