Fetching the paper…
Reading the bibliography…
Mixture-of-Experts (MoE) models improve the efficiency and scalability of dense language models by routing each token to a small number of experts in each layer.
Adaptive mixtures of local experts
R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the em algorithm
M. I. Jordan and R. A. Jacobs · 1994
Earlier work this paper cites.
Security engineering: a guide to building dependable distributed systems
R. J. Anderson · 2010
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer, 2017
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean · 2017
Earlier work this paper cites.
Make topk sort stable, 2019
Issues · 2019
Earlier work this paper cites.
Indices returned by torch.topk is of wrong order, 2020
Issues · 2020
Earlier work this paper cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
D. Lepikhin, H. Lee, Y. Xu, D. Chen, O. Firat, Y. Huang, M. Krikun, N. Shazeer, and Z. Chen · 2020
Earlier work this paper cites.
Scaling vision with sparse mixture of experts, 2021
C. Riquelme, J. Puigcerver, B. Mustafa, M. Neumann, R. Jenatton, A. S. Pinto, D. Keysers, and N. Houlsby · 2021
Earlier work this paper cites.
Glam: Efficient scaling of language models with mixture-of-experts, 2022
N. Du, Y. Huang, A. M. Dai, S. Tong, D. Lepikhin, Y. Xu, M. Krikun, Y. Zhou, A. W. Yu, O. Firat, B. Zoph, L. Fedus, M. Bosma, Z. Zhou, T. Wang, Y. E. Wang, K. Webster, M. Pellat, K. Robinson, K. Meier-Hellstern, T. Duke, L. Dixon, K. Zhang, Q. V. Le, Y. Wu, Z. Chen, and C. Cui · 2022
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2022
W. Fedus, B. Zoph, and N. Shazeer · 2022
Cited alongside, same era.
Learning to merge tokens in vision transformers, 2022
C. Renggli, A. S. Pinto, N. Houlsby, B. Mustafa, J. Puigcerver, and C. Riquelme · 2022
Cited alongside, same era.
Scaling up models and data with t5x
A. Roberts, H. W. Chung, A. Levskaya, G. Mishra, J. Bradbury, D. Andor, S. Narang, B. Lester, C. Gaffney, A. Mohiuddin, C. Hawthorne, A. Lewkowycz, A. Salcianu, M. van Zee, J. Austin, S. Goodman, L. B. Soares, H. Hu, S. Tsvyashchenko, A. Chowdhery, J. Bastings, J. Bulian, X. Garcia, J. Ni, A. Chen, K. Kenealy, J. H. Clark, S. Lee, D. Garrette, J. Lee-Thorp, C. Raffel, N. Shazeer, M. Ritter, M. Bosma, A. Passos, J. Maitin-Shepard, N. Fiedel, M. Omernick, B. Saeta, R. Sepassi, A. Spiridonov, J. Newlan, and A. Gesmundo · 2022
A survey on mixture of experts, 2024
W. Cai, J. Jiang, F. Wang, J. Tang, S. Kim, and J. Huang · 2024
Closest in time.
Buffer overflow in mixture of experts, 2024
J. Hayes, I. Shumailov, and I. Yona · 2024
Closest in time.
torch.topk returning unexpected output, 2024
Issues · 2024
Closest in time.
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. de las Casas, E. B. Hanna, F. Bressand, G. Lengyel, G. Bour, G. Lample, L. R. Lavaud, L. Saulnier, M.-A. Lachaux, P. Stock, S. Subramanian, S. Yang, S. Antoniak, T. L. Scao, T. Gervet, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed · 2024
Closest in time.
Pytorch cuda implementation of sort, which is used in topk, 2024
PyTorch · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
St-moe: Designing stable and transferable sparse expert models, 2022
B. Zoph, I. Bello, S. Kumar, N. Du, Y. Huang, J. Dean, N. Shazeer, and W. Fedus · 2022
Cited alongside, same era.
Privacy side channels in machine learning systems, 2023
E. Debenedetti, G. Severi, N. Carlini, C. A. Choquette-Choo, M. Jagielski, M. Nasr, E. Wallace, and F. Tramèr · 2023
Cited alongside, same era.
Tutel: Adaptive mixture-of-experts at scale
C. Hwang, W. Cui, Y. Xiong, Z. Yang, Z. Liu, H. Hu, Z. Wang, R. Salas, J. Jose, P. Ram, et al · 2023
Cited alongside, same era.
Grok-1, 2023
xAI · 2023
Cited alongside, same era.
Prompt stealing attacks against text-to-image generation models, 2024
X. Shen, Y. Qu, M. Backes, and Y. Zhang · 2024
Closest in time.
Gemini: A family of highly capable multimodal models, 2024
G. Team · 2024
Closest in time.
Mixture-of-experts with expert choice routing
Y. Zhou, T. Lei, H. Liu, N. Du, Y. Huang, V. Y. Zhao, A. Dai, Z. Chen, Q. Le, and J. Laudon · 2024
Closest in time.