Fetching the paper…
Reading the bibliography…
The cosine router in Mixture of Experts (MoE) has recently emerged as an attractive alternative to the conventional linear router.
Perplexity—a measure of the difficulty of speech recognition tasks
Fred Jelinek, Robert L Mercer, Lalit R Bahl, and James K Baker · 1977
Earlier work this paper cites.
A maximum likelihood approach to continuous speech recognition
Lalit R. Bahl, Frederick Jelinek, and Robert L. Mercer · 1983
Earlier work this paper cites.
Adaptive mixtures of local experts
Robert A. Jacobs, Michael I. Jordan, Steven J. Nowlan, and Geoffrey E. Hinton · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the EM algorithm
M. I. Jordan and R. A. Jacobs · 1994
Earlier work this paper cites.
Mixture models: Theory, geometry and applications
Bruce G. Lindsay · 1995
Earlier work this paper cites.
Bayesian Inference in Mixtures-of-Experts and Hierarchical Mixtures-of-Experts Models With an Application to Speech Recognition
Fengchun Peng, Robert A. Jacobs, and Martin A. Tanner · 1996
Earlier work this paper cites.
Assouad, Fano, and Le Cam
Bin Yu · 1997
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, and Pascal Vincent · 2000
Earlier work this paper cites.
Empirical processes in M-estimation
Sara van de Geer · 2000
Earlier work this paper cites.
Large text compression benchmark
Matt Mahoney · 2011
Earlier work this paper cites.
On convergence rates of mixtures of polynomial experts
Eduardo F. Mendes and Wenxin Jiang · 2012
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Alex Graves · 2013
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2017
Diederik P. Kingma and Jimmy Ba · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
On the properties of the softmax function with application in game theory and reinforcement learning
Bolin Gao and Lacra Pavel · 2018
Earlier work this paper cites.
Strong identifiability and optimal minimax rates for finite mixture estimation
P. Heinrich and J. Kahn · 2018
Cited alongside, same era.
Conformer: Convolution-augmented Transformer for Speech Recognition
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam M. Shazeer · 2021
Cited alongside, same era.
In search of lost domain generalization
Ishaan Gulrajani and David Lopez-Paz · 2021
Cited alongside, same era.
Domain generalization: A survey
Kaiyang Zhou, Ziwei Liu, Yu Qiao, Tao Xiang, and Chen Change Loy · 2022
Later among the works it cites.
A Mixture-of-Expert Approach to RL-based Dialogue Management
Yinlam Chow, Azamat Tulepbergenov, Ofir Nachum, Dhawal Gupta, Moonkyung Ryu, Mohammad Ghavamzadeh, and Craig Boutilier · 2023
Later among the works it cites.
Minimax optimal rate for parameter estimation in multivariate deviated models
Dat Do, Huy Nguyen, Khai Nguyen, and Nhat Ho · 2023
Later among the works it cites.
Improving expert specialization in mixture of experts
Yamuna Krishnamurthy, Chris Watkins, and Thomas Gaertner · 2023
Later among the works it cites.
Sparse mixture-of-experts are domain generalizable learners
Bo Li, Yifei Shen, Jingkang Yang, Yezhen Wang, Jiawei Ren, Tong Che, Jun Zhang, and Ziwei Liu · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hussein Hazimeh, Zhe Zhao, Aakanksha Chowdhery, Maheswaran Sathiamoorthy, Yihua Chen, Rahul Mazumder, Lichan Hong, and Ed Chi · 2021
Cited alongside, same era.
Non-asymptotic model selection in block-diagonal mixture of polynomial experts models
TrungTin Nguyen, Faicel Chamroukhi, Hien Duy Nguyen, and Florence Forbes · 2021
Cited alongside, same era.
Scaling vision with sparse mixture of experts
Carlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann, Rodolphe Jenatton, André Susano Pinto, Daniel Keysers, and Neil Houlsby · 2021
Cited alongside, same era.
Speechmoe: Scaling to large acoustic models with dynamic routing mixture of experts
Zhao You, Shulin Feng, Dan Su, and Dong Yu · 2021
Cited alongside, same era.
Towards understanding the mixture-of-experts layer in deep learning
Zixiang Chen, Yihe Deng, Yue Wu, Quanquan Gu, and Yuanzhi Li · 2022
Cited alongside, same era.
On the representation collapse of sparse mixture of experts
Zewen Chi, Li Dong, Shaohan Huang, Damai Dai, Shuming Ma, Barun Patra, Saksham Singhal, Payal Bajaj, Xia Song, Xian-Ling Mao, Heyan Huang, and Furu Wei · 2022
Cited alongside, same era.
Spatial mixture-of-experts
Nikoli Dryden and Torsten Hoefler · 2022
Cited alongside, same era.
Huy Nguyen, TrungTin Nguyen, and Nhat Ho · 2023
Later among the works it cites.
Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models
Damai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y. Wu, Zhenda Xie, Y. K. Li, Panpan Huang, Fuli Luo, Chong Ruan, Zhifang Sui, and Wenfeng Liang · 2024
Closest in time.
Fusemoe: Mixture-of-experts transformers for fleximodal fusion
Xing Han, Huy Nguyen, Carl Harris, Nhat Ho, and Suchi Saria · 2024
Closest in time.
Mixtral of experts
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2024
Closest in time.
Mixture of experts meets prompt-based continual learning
Minh Le, An Nguyen, Huy Nguyen, Trang Nguyen, Trang Pham, Linh Van Ngo, and Nhat Ho · 2024
Closest in time.
Competesmoe – effective training of sparse mixture of experts via competition
Quang Pham, Giang Do, Huy Nguyen, TrungTin Nguyen, Chenghao Liu, Mina Sartipi, Binh T. Nguyen, Savitha Ramasamy, Xiaoli Li, Steven Hoi, and Nhat Ho · 2024
Closest in time.
From sparse to soft mixtures of experts
Joan Puigcerver, Carlosx Riquelme, Basil Mustafa, and Neil Houlsby · 2024
Closest in time.
Revisiting prefix-tuning: Statistical benefits of reparameterization among prompts
Minh Le, Chau Nguyen, Huy Nguyen, Quyen Tran, Trung Le, and Nhat Ho · 2025
Closest in time.
Theory on mixture-of-experts in continual learning
Hongbo Li, Sen Lin, Lingjie Duan, Yingbin Liang, and Ness Shroff · 2025
Closest in time.
Understanding expert structures on minimax parameter estimation in contaminated mixture of experts
Fanqi Yan, Huy Nguyen, Dung Le, Pedram Akbarian, and Nhat Ho · 2025
Closest in time.