Fetching the paper…
Reading the bibliography…
The Mixture of Experts (MoE) has emerged as a highly successful technique in deep learning, based on the principle of divide-and-conquer to maximize model capacity without significant additional computational cost.
Roberta: A robustly optimized bert pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 1907
Earlier work this paper cites.
Adaptive mixtures of local experts
R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton · 1991
Earlier work this paper cites.
Generalized inverses: theory and applications , volume 15
A. Ben-Israel and T. N. Greville · 2003
Earlier work this paper cites.
Introduction to the conll-2003 shared task: Language-independent named entity recognition
E. F. Sang and F. De Meulder · 2003
Earlier work this paper cites.
Adaptive filter theory. pearson education india
S. Haykin · 2008
Earlier work this paper cites.
Projection matrices, generalized inverse matrices, and singular value decomposition
H. Yanai, K. Takeuchi, and Y. Takane · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang · 2016
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. V. Le, G. E. Hinton, and J. Dean · 2017
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
The University of Sydney’s machine translation system for WMT19
L. Ding and D. Tao · 2019
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2019
Earlier work this paper cites.
Mixture models for diverse machine translation: Tricks of the trade
T. Shen, M. Ott, M. Auli, and M. Ranzato · 2019
Cited alongside, same era.
Continual learning of context-dependent processing in neural networks
G. Zeng, Y. Chen, B. Cui, and S. Yu · 2019
Cited alongside, same era.
ELECTRA: pre-training text encoders as discriminators rather than generators
K. Clark, M. Luong, Q. V. Le, and C. D. Manning · 2020
Cited alongside, same era.
ALBERT: A lite BERT for self-supervised learning of language representations
Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut · 2020
Cited alongside, same era.
Understanding and improving lexical choice in non-autoregressive translation
L. Ding, L. Wang, X. Liu, D. F. Wong, D. Tao, and Z. Tu · 2021
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2021
Glam: Efficient scaling of language models with mixture-of-experts
N. Du, Y. Huang, A. M. Dai, S. Tong, D. Lepikhin, Y. Xu, M. Krikun, Y. Zhou, A. W. Yu, O. Firat, et al · 2022
Later among the works it cites.
Parameter-efficient mixture-of-experts architecture for pre-trained language models
Z. Gao, P. Liu, W. X. Zhao, Z. Lu, and J. Wen · 2022
Later among the works it cites.
Sparsely activated mixture-of-experts are robust multi-task learners
S. Gupta, S. Mukherjee, K. Subudhi, E. Gonzalez, D. Jose, A. H. Awadallah, and J. Gao · 2022
Later among the works it cites.
Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale
S. Rajbhandari, C. Li, Z. Yao, M. Zhang, R. Y. Aminabadi, A. A. Awan, J. Rasley, and Y. He · 2022
Later among the works it cites.
A contrastive cross-channel data augmentation framework for aspect-based sentiment analysis
B. Wang, L. Ding, Q. Zhong, X. Li, and D. Tao · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
W. Fedus, B. Zoph, and N. Shazeer · 2021
Cited alongside, same era.
Gshard: Scaling giant models with conditional computation and automatic sharding
D. Lepikhin, H. Lee, Y. Xu, D. Chen, O. Firat, Y. Huang, M. Krikun, N. Shazeer, and Z. Chen · 2021
Cited alongside, same era.
Base layers: Simplifying training of large, sparse models
M. Lewis, S. Bhosale, T. Dettmers, N. Goyal, and L. Zettlemoyer · 2021
Cited alongside, same era.
Scaling vision with sparse mixture of experts
C. Riquelme, J. Puigcerver, B. Mustafa, M. Neumann, R. Jenatton, A. Susano Pinto, D. Keysers, and N. Houlsby · 2021
Cited alongside, same era.
Hash layers for large sparse models
S. Roller, S. Sukhbaatar, J. Weston, et al · 2021
Cited alongside, same era.
M6-t: Exploring sparse expert models and beyond
A. Yang, J. Lin, R. Men, C. Zhou, L. Jiang, X. Jia, A. Wang, J. Zhang, J. Wang, Y. Li, et al · 2021
Cited alongside, same era.
StableMoE: Stable routing strategy for mixture of experts
D. Dai, L. Dong, S. Ma, B. Zheng, Z. Sui, B. Chang, and F. Wei · 2022
Cited alongside, same era.
Eliciting transferability in multi-task learning with task-level mixture-of-experts
Q. Ye, J. Zha, and X. Ren · 2022
Later among the works it cites.
Bort: Towards explainable neural networks with bounded orthogonal constraint
B. Zhang, W. Zheng, J. Zhou, and J. Lu · 2022
Later among the works it cites.
Mixture-of-experts with expert choice routing
Y. Zhou, T. Lei, H. Liu, N. Du, Y. Huang, V. Zhao, A. Dai, Z. Chen, Q. Le, and J. Laudon · 2022
Later among the works it cites.
Taming sparsely activated transformer with stochastic experts
S. Zuo, X. Liu, J. Jiao, Y. J. Kim, H. Hassan, R. Zhang, J. Gao, and T. Zhao · 2022
Later among the works it cites.
Improved training of mixture-of-experts language gans
Y. Chai, Q. Yin, and J. Zhang · 2023
Closest in time.
Merging experts into one: Improving computational efficiency of mixture of experts
S. He, R.-Z. Fan, L. Ding, L. Shen, T. Zhou, and D. Tao · 2023
Closest in time.
Dynamic regularized sharpness aware minimization in federated learning: Approaching global consistency and smooth landscape
Y. Sun, L. Shen, S.-Y. Chen, L. Ding, and D. Tao · 2023
Closest in time.
Divide, conquer, and combine: Mixture of semantic-independent experts for zero-shot dialogue state tracking
Q. Wang, L. Ding, Y. Cao, Y. Zhan, Z. Lin, S. Wang, D. Tao, and L. Guo · 2023
Closest in time.