Fetching the paper…
Reading the bibliography…
Sparse Mixture-of-Experts (SMoE) models represent a significant advancement in large language model (LLM) development through their efficient parameter utilization.
Fcm: The fuzzy c-means clustering algorithm
Bezdek, J. C., Ehrlich, R., and Full, W · 1984
Earlier work this paper cites.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
The fifth pascal recognizing textual entailment challenge
Bentivogli, L., Clark, P., Dagan, I., and Giampiccolo, D · 2009
Earlier work this paper cites.
Convergent learning: Do different neural networks learn the same representations?
Li, Y., Yosinski, J., Clune, J., Lipson, H., and Hopcroft, J. E · 2016
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Sigmoid-weighted linear units for neural network function approximation in reinforcement learning
Elfwing, S., Uchibe, E., and Doya, K · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A · 2018
Earlier work this paper cites.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Codeqa: A question answering dataset for source code comprehension
Liu, C. and Wan, X · 2021
Cited alongside, same era.
Winogrande: an adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2021
Cited alongside, same era.
On the opportunities and risks of foundation models, 2022
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., et al · 2022
Cited alongside, same era.
Task-specific expert pruning for sparse mixture-of-experts
Chen, T., Huang, S., Xie, Y., Jiao, B., Jiang, D., Zhou, H., Li, J., and Wei, F · 2022
Cited alongside, same era.
Palm: Scaling language modeling with pathways, 2022
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., et al · 2022
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
TIES-merging: Resolving interference when merging models
Yadav, P., Tam, D., Choshen, L., Raffel, C., and Bansal, M · 2023
Later among the works it cites.
A framework for few-shot language model evaluation, 2024
Gao, L., Tow, J., Abbasi, B., Biderman, S., Black, S., et al · 2024
Closest in time.
Demystifying the compression of mixture-of-experts through a unified framework
He, S., Dong, D., Ding, L., and Li, A · 2024
Closest in time.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., et al · 2024
Closest in time.
Merge, then compress: Demystify efficient SMoe with hints from its routing policy
Li, P., Zhang, Z., Yadav, P., Sung, Y.-L., Cheng, Y., Bansal, M., and Chen, T · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fedus, W., Zoph, B., and Shazeer, N · 2022
Cited alongside, same era.
K-means clustering algorithms: A comprehensive review, variants analysis, and advances in the era of big data
Ikotun, A. M., Ezugwu, A. E., Abualigah, L., Abuhaija, B., and Heming, J · 2022
Cited alongside, same era.
Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering
Pal, A., Umapathi, L. K., and Sankarasubbu, M · 2022
Cited alongside, same era.
Diversifying the mixture-of-experts representation for language models with orthogonal optimizer
Liu, B., Ding, L., Shen, L., Peng, K., Cao, Y., Cheng, D., and Tao, D · 2023
Cited alongside, same era.
Approximation bounds for hierarchical clustering: Average linkage, bisecting k-means, and local search
Moseley, B. and Wang, J. R · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Cited alongside, same era.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J
Cited in the paper.
Not all experts are equal: Efficient expert pruning and skipping for mixture-of-experts large language models
Lu, X., Liu, Q., Xu, Y., Zhou, A., Huang, S., Zhang, B., Yan, J., and Li, H · 2024
Closest in time.
GPT-4 technical report, 2024
OpenAI, Achiam, J., Adler, S., Agarwal, S., Ahmad, L., et al · 2024
Closest in time.
Zipit! merging models from different tasks without training
Stoica, G., Bolya, D., Bjorner, J. B., Ramesh, P., Hearn, T., and Hoffman, J · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024
Team, G., Georgiev, P., Lei, V. I., Burnell, R., Bai, L., and others · 2024
Closest in time.
Qwen1.5-moe: Matching 7b model performance with 1/3 activated parameters”, February 2024
Team, Q · 2024
Closest in time.
Training-free pretrained model merging
Xu, Z., Yuan, K., Wang, H., Wang, Y., Song, M., and Song, J · 2024
Closest in time.