Fetching the paper…
Reading the bibliography…
In the past year, large (>100B parameter) mixture-of-expert (MoE) models have become increasingly common in the open domain.
Adaptive mixtures of local experts
Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton. 1991 · 1991
Earlier work this paper cites.
Conditional routing of information to the cortex: a model of the basal ganglia’s role in cognitive coordination
Andrea Stocco, Christian Lebiere, and John R Anderson. 2010 · 2010
Earlier work this paper cites.
Wic: the word-in-context dataset for evaluating context-sensitive meaning representations
Mohammad Taher Pilehvar and Jose Camacho-Collados. 2018 · 2018
Earlier work this paper cites.
Swords: A benchmark for lexical substitution with improved data coverage and quality
Mina Lee, Chris Donahue, Robin Jia, Alexander Iyabor, and Percy Liang. 2021 · 2021
Earlier work this paper cites.
Why and how the brain weights contributions from a mixture of experts
John P O’Doherty, Sang Wan Lee, Reza Tadayonnejad, Jeff Cockburn, Kyo Iigaya, and Caroline J Charpentier. 2021 · 2021
Earlier work this paper cites.
Hash layers for large sparse models
Stephen Roller, Sainbayar Sukhbaatar, Jason Weston, et al. 2021 · 2021
Earlier work this paper cites.
Taming sparsely activated transformer with stochastic experts
Simiao Zuo, Xiaodong Liu, Jian Jiao, Young Jin Kim, Hany Hassan, Ruofei Zhang, Tuo Zhao, and Jianfeng Gao. 2021 · 2021
Earlier work this paper cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer. 2022 · 2022
Earlier work this paper cites.
Does bert rediscover a classical nlp pipeline?
Jingcheng Niu, Wenjie Lu, and Gerald Penn. 2022 · 2022
Cited alongside, same era.
St-moe: Designing stable and transferable sparse expert models
Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus. 2022 · 2022
Cited alongside, same era.
Sparse autoencoders find highly interpretable features in language models
Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey. 2023 · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023 · 2023
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. 2024 · 2024
Later among the works it cites.
From tokens to words: On the inner lexicon of llms
Guy Kaplan, Matanel Oren, Yuval Reif, and Roy Schwartz. 2024 · 2024
Later among the works it cites.
Jamba: A hybrid transformer-mamba language model
Opher Lieber, Barak Lenz, Hofit Bata, Gal Cohen, Jhonathan Osin, Itay Dalmedigos, Erez Safahi, Shaked Meirom, Yonatan Belinkov, Shai Shalev-Shwartz, et al. 2024 · 2024
Later among the works it cites.
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xiaozhe Ren, Pingyi Zhou, Xinfan Meng, Xinjing Huang, Yadao Wang, Weichao Wang, Pengfei Li, Xiaoda Zhang, Alexander Podolskiy, Grigory Arshinov, et al. 2023 · 2023
Cited alongside, same era.
Towards an empirical understanding of moe design choices
Dongyang Fan, Bettina Messmer, and Martin Jaggi. 2024 · 2024
Cited alongside, same era.
Peter Jansen, Marc-Alexandre Côté, Tushar Khot, Erin Bransom, Bhavana Dalvi Mishra, Bodhisattwa Prasad Majumder, Oyvind Tafjord, and Peter Clark. 2024 · 2024
Cited alongside, same era.
Fuzhao Xue, Zian Zheng, Yao Fu, Jinjie Ni, Zangwei Zheng, Wangchunshu Zhou, and Yang You. 2024 · 2024
Later among the works it cites.
Llama 4: Multimodal and multilingual mixture-of-experts foundation models
Meta AI. 2025 · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. 2025 · 2025
Closest in time.