Fetching the paper…
Reading the bibliography…
The softmax gating function is arguably the most popular choice in mixture of experts modeling.
Adaptive mixtures of local experts
R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton · 1991
Earlier work this paper cites.
Universal approximation bounds for superpositions of a sigmoidal function
A. Barron · 1993
Earlier work this paper cites.
Hierarchical mixtures of experts and the EM algorithm
M. I. Jordan and R. A. Jacobs · 1994
Earlier work this paper cites.
Bayesian Inference in Mixtures-of-Experts and Hierarchical Mixtures-of-Experts Models With an Application to Speech Recognition
F. Peng, R. A. Jacobs, and M. A. Tanner · 1996
Earlier work this paper cites.
Assouad, Fano, and Le Cam
B. Yu · 1997
Earlier work this paper cites.
Empirical processes in M-estimation
S. van de Geer · 2000
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Conformer: Convolution-augmented Transformer for Speech Recognition
A. Gulati, J. Qin, C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang · 2020
Earlier work this paper cites.
DSelect-k: Differentiable Selection in the Mixture of Experts with Applications to Multi-Task Learning
H. Hazimeh, Z. Zhao, A. Chowdhery, M. Sathiamoorthy, Y. Chen, R. Mazumder, L. Hong, and E. Chi · 2021
Earlier work this paper cites.
GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
D. Lepikhin, H. Lee, Y. Xu, D. Chen, O. Firat, Y. Huang, M. Krikun, N. Shazeer, and Z. Chen · 2021
Earlier work this paper cites.
Scaling vision with sparse mixture of experts
C. Riquelme, J. Puigcerver, B. Mustafa, M. Neumann, R. Jenatton, A. S. Pint, D. Keysers, and N. Houlsby · 2021
Earlier work this paper cites.
Scaling vision with sparse mixture of experts
C. Ruiz, J. Puigcerver, B. Mustafa, M. Neumann, R. Jenatton, A. Pinto, D. Keysers, and N. Houlsby · 2021
Earlier work this paper cites.
Speechmoe: Scaling to large acoustic models with dynamic routing mixture of experts
Z. You, S. Feng, D. Su, and D. Yu · 2021
Earlier work this paper cites.
On the representation collapse of sparse mixture of experts
Z. Chi, L. Dong, S. Huang, D. Dai, S. Ma, B. Patra, S. Singhal, P. Bajaj, X. Song, X.-L. Mao, H. Huang, and F. Wei · 2022
Cited alongside, same era.
Glam: Efficient scaling of language models with mixture-of-experts
N. Du, Y. Huang, A. M. Dai, S. Tong, D. Lepikhin, Y. Xu, M. Krikun, Y. Zhou, A. Yu, O. Firat, B. Zoph, L. Fedus, M. Bosma, Z. Zhou, T. Wang, E. Wang, K. Webster, M. Pellat, K. Robinson, K. Meier-Hellstern, T. Duke, L. Dixon, K. Zhang, Q. Le, Y. Wu, Z. Chen, and C. Cui · 2022
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
W. Fedus, B. Zoph, and N. Shazeer · 2022
Cited alongside, same era.
Sparsely activated mixture-of-experts are robust multi-task learners
S. Gupta, S. Mukherjee, K. Subudhi, E. Gonzalez, D. Jose, A. Awadallah, and J. Gao · 2022
Cited alongside, same era.
Convergence rates for Gaussian mixtures of experts
N. Ho, C.-Y. Yang, and M. I. Jordan · 2022
Cited alongside, same era.
Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models
D. Dai, C. Deng, C. Zhao, R. X. Xu, H. Gao, D. Chen, J. Li, W. Zeng, X. Yu, Y. Wu, Z. Xie, Y. K. Li, P. Huang, F. Luo, C. Ruan, Z. Sui, and W. Liang · 2024
Closest in time.
Fusemoe: Mixture-of-experts transformers for fleximodal fusion
X. Han, H. Nguyen, C. Harris, N. Ho, and S. Saria · 2024
Closest in time.
X. O. He · 2024
Closest in time.
Mixtral of experts
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. de las Casas, E. B. Hanna, F. Bressand, G. Lengyel, G. Bour, G. Lample, L. R. Lavaud, L. Saulnier, M.-A. Lachaux, P. Stock, S. Subramanian, S. Yang, S. Antoniak, T. L. Scao, T. Gervet, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed · 2024
Closest in time.
Is temperature sample efficient for softmax Gaussian mixture of experts?
H. Nguyen, P. Akbarian, and N. Ho · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sparse mixers: Combining moe and mixing to build a more efficient BERT
J. Lee-Thorp and J. Ainslie · 2022
Cited alongside, same era.
M 3 ViT: Mixture-of-Experts Vision Transformer for Efficient Multi-task Learning with Model-Accelerator Co-design
H. Liang, Z. Fan, R. Sarkar, Z. Jiang, T. Chen, K. Zou, Y. Cheng, C. Hao, and Z. Wang · 2022
Cited alongside, same era.
Refined convergence rates for maximum likelihood estimation under finite mixture models
T. Manole and N. Ho · 2022
Cited alongside, same era.
A Mixture-of-Expert Approach to RL-based Dialogue Management
Y. Chow, A. Tulepbergenov, O. Nachum, D. Gupta, M. Ryu, M. Ghavamzadeh, and C. Boutilier · 2023
Cited alongside, same era.
Approximating two-layer feedforward networks for efficient transformers
R. Csordás, K. Irie, and J. Schmidhuber · 2023
Cited alongside, same era.
Minimax optimal rate for parameter estimation in multivariate deviated models
D. Do, H. Nguyen, K. Nguyen, and N. Ho · 2023
Cited alongside, same era.
Demystifying softmax gating function in Gaussian mixture of experts
H. Nguyen, T. Nguyen, and N. Ho · 2023
Cited alongside, same era.
A general theory for softmax gating multinomial logistic mixture of experts
H. Nguyen, P. Akbarian, T. Nguyen, and N. Ho · 2024
Closest in time.
Statistical advantages of perturbing cosine router in mixture of experts
H. Nguyen, P. Akbarian, T. Pham, T. Nguyen, S. Zhang, and N. Ho · 2024
Closest in time.
Statistical perspective of top-k sparse softmax gating mixture of experts
H. Nguyen, P. Akbarian, F. Yan, and N. Ho · 2024
Closest in time.
On least square estimation in softmax gating mixture of experts
H. Nguyen, N. Ho, and A. Rinaldo · 2024
Closest in time.
Mixtures of experts unlock parameter scaling for deep RL
J. S. Obando Ceron, G. Sokar, T. Willi, C. Lyle, J. Farebrother, J. N. Foerster, G. K. Dziugaite, D. Precup, and P. S. Castro · 2024
Closest in time.
Competesmoe – effective training of sparse mixture of experts via competition
Q. Pham, G. Do, H. Nguyen, T. Nguyen, C. Liu, M. Sartipi, B. T. Nguyen, S. Ramasamy, X. Li, S. Hoi, and N. Ho · 2024
Closest in time.
From sparse to soft mixtures of experts
J. Puigcerver, C. Riquelme, B. Mustafa, and N. Houlsby · 2024
Closest in time.
Understanding expert structures on minimax parameter estimation in contaminated mixture of experts
F. Yan, H. Nguyen, D. Le, P. Akbarian, and N. Ho · 2024
Closest in time.