Fetching the paper…
Reading the bibliography…
Artificial intelligence (AI) has achieved astonishing successes in many domains, especially with the recent breakthroughs in the development of foundational large models.
S. Kariyappa and M. Qureshi, “Improving adversarial robustness of ensembles with diversity training. 2019,” Available: arXiv
1901
Earlier work this paper cites.
J. Ma, Z. Zhao, X. Yi, J. Chen, L. Hong, and E. H. Chi, “Modeling task relationships in multi-task learning with multi-gate mixture-of-experts,” in Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining
1939
Earlier work this paper cites.
M. Aitkin and N. Longford, “Statistical modelling issues in school effectiveness studies,” Journal of the Royal Statistical Society: Series A (General)
1986
Earlier work this paper cites.
H. Goldstein, “Multilevel mixed linear model analysis using iterative generalized least squares,” Biometrika
1986
Earlier work this paper cites.
M. McCloskey and N. J. Cohen, “Catastrophic interference in connectionist networks: The sequential learning problem,” in Psychology of learning and motivation
1989
Earlier work this paper cites.
P. F. Brown, J. Cocke, S. A. Della Pietra, V. J. Della Pietra, F. Jelinek, J. Lafferty, R. L. Mercer, and P. S. Roossin, “A statistical approach to machine translation,” Computational linguistics
1990
Earlier work this paper cites.
R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton, “Adaptive mixtures of local experts,” Neural computation
1991
Earlier work this paper cites.
M. P. Keane, “A note on identification in the multinomial probit model,” Journal of Business & Economic Statistics
1992
Earlier work this paper cites.
M. I. Jordan and R. A. Jacobs, “Hierarchical mixtures of experts and the em algorithm,” Neural computation
1994
Earlier work this paper cites.
L. Xu, M. Jordan, and G. E. Hinton, “An alternative model for mixtures of experts,” Advances in neural information processing systems
1994
Earlier work this paper cites.
J. Geweke, M. Keane, and D. Runkle, “Alternative computational approaches to inference in the multinomial probit model,” The review of economics and statistics
1994
Earlier work this paper cites.
R. Caruana, “Multitask learning,” Machine learning
1997
Earlier work this paper cites.
E. Reiter and R. Dale, “Building applied natural language generation systems,” Natural Language Engineering
1997
Earlier work this paper cites.
A. J. Zeevi, R. Meir, and V. Maiorov, “Error bounds for functional approximation and estimation using mixtures of experts,” IEEE Transactions on Information Theory
1998
Earlier work this paper cites.
W. Jiang and M. A. Tanner, “Hierarchical mixtures-of-experts for exponential family regression models: approximation and maximum likelihood estimation,” Annals of Statistics
1999
Earlier work this paper cites.
W. Jiang and M. A. Tanner, “On the approximation rate of hierarchical mixtures-of-experts for generalized linear models,” Neural computation
1999
Earlier work this paper cites.
MIT press, 1999
C. Manning and H. Schutze, Foundations of statistical natural language processing · 1999
Earlier work this paper cites.
Pearson Education India, 2000
D. Jurafsky, Speech & language processing · 2000
Earlier work this paper cites.
K. Doya, K. Samejima, K.-i. Katagiri, and M. Kawato, “Multiple model-based reinforcement learning,” Neural computation
2002
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the 40th annual meeting of the Association for Computational Linguistics
2002
Earlier work this paper cites.
K. Samejima, K. Doya, and M. Kawato, “Inter-module credit assignment in modular reinforcement learning,” Neural Networks
2003
Earlier work this paper cites.
Y. Takahashi and M. Asada, “Modular learning systems for soccer robot,” in Proceedings of the Fourth International Symposium on Human and Artificial Intelligence Systems
2004
Earlier work this paper cites.
Y. Takahashi, K. Edazawa, and M. Asada, “Modular learning system and scheduling for behavior acquisition in multi-agent environment,” in RoboCup 2004: Robot Soccer World Cup VIII 8
2005
Earlier work this paper cites.
Y. Takahashi, K. Edazawa, K. Noma, and M. Asada, “Simultaneous learning to acquire competitive behaviors in multi-agent system based on a modular learning system,” in 2005 IEEE/RSJ International Conference on Intelligent Robots and Systems
2005
Earlier work this paper cites.
J. Geweke and M. Keane, “Smoothly mixing regressions,” Journal of Econometrics
2007
Earlier work this paper cites.
H. Van Seijen, B. Bakker, L. Kester, et al
2008
Earlier work this paper cites.
John Wiley & Sons, 2011
H. Goldstein, Multilevel statistical models · 2011
Earlier work this paper cites.
E. F. Mendes and W. Jiang, “On convergence rates of mixtures of polynomial experts,” Neural computation
2012
Earlier work this paper cites.
S. E. Yuksel, J. N. Wilson, and P. D. Gader, “Twenty years of mixture of experts,” IEEE transactions on neural networks and learning systems
2012
Earlier work this paper cites.
S. Ingrassia, S. C. Minotti, and G. Vittadini, “Local statistical modeling via a cluster-weighted approach with elliptical distributions,” Journal of classification
2012
Earlier work this paper cites.
A. Graves and A. Graves, “Long short-term memory,” Supervised sequence labelling with recurrent neural networks
2012
Earlier work this paper cites.
I. Szita, “Reinforcement learning in games,” in Reinforcement Learning: State-of-the-art
2012
Earlier work this paper cites.
K. Mülling, J. Kober, O. Kroemer, and J. Peters, “Learning to select and generalize striking movements in robot table tennis,” The International Journal of Robotics Research
2013
Earlier work this paper cites.
J. Kober, J. A. Bagnell, and J. Peters, “Reinforcement learning in robotics: A survey,” The International Journal of Robotics Research
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
S. Masoudnia and R. Ebrahimpour, “Mixture of experts: a literature survey,” Artificial Intelligence Review
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
M. Ewerton, G. Neumann, R. Lioutikov, H. B. Amor, J. Peters, and G. Maeda, “Learning multiple collaborative tasks with a mixture of interaction primitives,” in 2015 IEEE International Conference on Robotics and Automation (ICRA)
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
J. Hirschberg and C. D. Manning, “Advances in natural language processing,” Science
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
X. B. Peng, G. Berseth, and M. Van de Panne, “Terrain-adaptive locomotion skills using deep reinforcement learning,” ACM Transactions on Graphics (TOG)
2016
Earlier work this paper cites.
P. Tommasino, D. Caligiore, M. Mirolli, and G. Baldassarre, “A reinforcement learning architecture that transfers knowledge between skills when solving multiple tasks,” IEEE Transactions on Cognitive and Developmental Systems
2016
Earlier work this paper cites.
H. D. Nguyen, L. R. Lloyd-Jones, and G. J. McLachlan, “A universal approximation theorem for mixture-of-experts models,” Neural computation
2016
Earlier work this paper cites.
A. Santoro, S. Bartunov, M. Botvinick, D. Wierstra, and T. Lillicrap, “Meta-learning with memory-augmented neural networks,” in International conference on machine learning
2016
Earlier work this paper cites.
O. Vinyals, C. Blundell, T. Lillicrap, D. Wierstra, et al
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Springer, 2016
D. Yu and L. Deng, Automatic speech recognition · 2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
C. Sun, A. Shrivastava, S. Singh, and A. Gupta, “Revisiting unreasonable effectiveness of data in deep learning era,” in Proceedings of the IEEE international conference on computer vision
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
S. Gross, M. Ranzato, and A. Szlam, “Hard mixtures of experts for large scale weakly supervised vision,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
2017
Earlier work this paper cites.
R. Aljundi, P. Chakravarty, and T. Tuytelaars, “Expert gate: Lifelong learning with a network of experts,” in Proceedings of the IEEE conference on computer vision and pattern recognition
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
V. C. Kumar, S. Ha, and C. K. Liu, “Learning a unified control policy for safe falling,” in 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems
2017
Earlier work this paper cites.
J. Snell, K. Swersky, and R. Zemel, “Prototypical networks for few-shot learning,” Advances in neural information processing systems
2017
Earlier work this paper cites.
C. Finn, P. Abbeel, and S. Levine, “Model-agnostic meta-learning for fast adaptation of deep networks,” in International conference on machine learning
2017
Earlier work this paper cites.
C. Devin, A. Gupta, T. Darrell, P. Abbeel, and S. Levine, “Learning modular neural network policies for multi-task and multi-robot transfer,” in 2017 IEEE international conference on robotics and automation (ICRA)
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas, “Communication-efficient learning of deep networks from decentralized data,” in Artificial intelligence and statistics
2017
Earlier work this paper cites.
J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska, et al
2017
Earlier work this paper cites.
B. Lakshminarayanan, A. Pritzel, and C. Blundell, “Simple and scalable predictive uncertainty estimation using deep ensembles,” Advances in neural information processing systems
2017
Earlier work this paper cites.
A. Kendall and Y. Gal, “What uncertainties do we need in bayesian deep learning for computer vision?,” Advances in neural information processing systems
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
C. Ledig, L. Theis, F. Huszár, J. Caballero, A. Cunningham, A. Acosta, A. Aitken, A. Tejani, J. Totz, Z. Wang, et al
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” in International conference on machine learning
2018
Earlier work this paper cites.
K.-H. Thung and C.-Y. Wee, “A brief review on multi-task learning,” Multimedia Tools and Applications
2018
Earlier work this paper cites.
Y. Zhang and Q. Yang, “An overview of multi-task learning,” National Science Review
2018
Earlier work this paper cites.
S. Karimpouli, S. Khoshlesan, E. H. Saenger, and H. H. Koochi, “Application of alternative digital rock physics methods in a real case study: a challenge between clean and cemented samples,” Geophysical Prospecting
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
L. Wei, S. Zhang, W. Gao, and Q. Tian, “Person transfer gan to bridge domain gap for person re-identification,” in Proceedings of the IEEE conference on computer vision and pattern recognition
2018
Earlier work this paper cites.
T. Young, D. Hazarika, S. Poria, and E. Cambria, “Recent trends in deep learning based natural language processing,” ieee Computational intelligenCe magazine
2018
Earlier work this paper cites.
L. Zhang, S. Wang, and B. Liu, “Deep learning for sentiment analysis: A survey,” Wiley interdisciplinary reviews: data mining and knowledge discovery
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers)
2019
Earlier work this paper cites.
Y. Yi, K.-Y. Chen, and H.-Y. Gu, “Mixture of CNN experts from multiple acoustic feature domain for music genre classification,” in 2019 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)
2019
Earlier work this paper cites.
L. Zhang, S. Huang, W. Liu, and D. Tao, “Learning a mixture of granularity-specific experts for fine-grained categorization,” in 2019 IEEE/CVF International Conference on Computer Vision (ICCV)
2019
Earlier work this paper cites.
X. B. Peng, M. Chang, G. Zhang, P. Abbeel, and S. Levine, “Mcp: Learning composable hierarchical control with multiplicative compositional policies,” Advances in neural information processing systems
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
R. Aljundi, K. Kelchtermans, and T. Tuytelaars, “Task-free continual learning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
2019
Earlier work this paper cites.
Q. Sun, Y. Liu, T.-S. Chua, and B. Schiele, “Meta-transfer learning for few-shot learning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
2019
Earlier work this paper cites.
K. Rakelly, A. Zhou, C. Finn, S. Levine, and D. Quillen, “Efficient off-policy meta-reinforcement learning via probabilistic context variables,” in International conference on machine learning
2019
Earlier work this paper cites.
N. Guha, A. Talwalkar, and V. Smith, “One-shot federated learning,” arXiv preprint arXiv:1902.11175
2019
Earlier work this paper cites.
X. Wang, Z. Cai, D. Gao, and N. Vasconcelos, “Towards universal object detection by domain attention,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
2019
Earlier work this paper cites.
T. Pang, K. Xu, C. Du, N. Chen, and J. Zhu, “Improving adversarial robustness via promoting ensemble diversity,” in International Conference on Machine Learning
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
C. Thompson Neil, G. Kristjan, L. Keeheon, and F. Manso Gabriel, “The computational limits of deep learning,” ArXiv, Cornell University, juillet
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Cited alongside, same era.
H. Tang, J. Liu, M. Zhao, and X. Gong, “Progressive layered extraction (ple): A novel multi-task learning (mtl) model for personalized recommendations,” in Proceedings of the 14th ACM conference on recommender systems
2020
Cited alongside, same era.
Y. Zhou, J. Gao, and T. Asfour, “Movement primitive learning and generalization: Using mixture density networks,” IEEE Robotics & Automation Magazine
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2023
Later among the works it cites.
C. E. Heinbaugh, E. Luz-Ricca, and H. Shao, “Data-free one-shot federated learning under very high statistical heterogeneity,” in The Eleventh International Conference on Learning Representations
2023
Later among the works it cites.
S. Su, B. Li, and X. Xue, “One-shot federated learning without server-side training,” Neural Networks
2023
Later among the works it cites.
M. N. R. Chowdhury, S. Zhang, M. Wang, S. Liu, and P.-Y. Chen, “Patch-level routing in mixture-of-experts is provably sample-efficient for convolutional neural networks,” in International Conference on Machine Learning
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Ghosh, J. Chung, D. Yin, and K. Ramchandran, “An efficient framework for clustered federated learning,” Advances in neural information processing systems
2020
Cited alongside, same era.
S. Pavlitskaya, C. Hubschneider, M. Weber, R. Moritz, F. Huger, P. Schlicht, and J. M. Zollner, “Using mixture of expert models to gain insights into semantic segmentation,” in 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)
2020
Cited alongside, same era.
X. Wang, F. Yu, L. Dunlap, Y.-A. Ma, R. Wang, A. Mirhoseini, T. Darrell, and J. E. Gonzalez, “Deep mixture of experts via shallow embedding,” in Uncertainty in artificial intelligence
2020
Cited alongside, same era.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of machine learning research
2020
Cited alongside, same era.
P.-E. Sarlin, D. DeTone, T. Malisiewicz, and A. Rabinovich, “Superglue: Learning feature matching with graph neural networks,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
2020
Cited alongside, same era.
2020
Cited alongside, same era.
N. Vithayathil Varghese and Q. H. Mahmoud, “A survey of multi-task deep reinforcement learning,” Electronics
2020
Cited alongside, same era.
R. Yang, H. Xu, Y. Wu, and X. Wang, “Multi-task reinforcement learning with soft modularization,” Advances in Neural Information Processing Systems
2020
Cited alongside, same era.
H. Nguyen, T. Nguyen, and N. Ho, “Demystifying softmax gating function in gaussian mixture of experts,” Advances in Neural Information Processing Systems
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Z. Xue, G. Song, Q. Guo, B. Liu, Z. Zong, Y. Liu, and P. Luo, “Raphael: Text-to-image generation via large mixture of diffusion paths,” Advances in Neural Information Processing Systems
2023
Later among the works it cites.
Y. Chai, Q. Yin, and J. Zhang, “Improved training of mixture-of-experts language gans,” in ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
W. Peebles and S. Xie, “Scalable diffusion models with transformers,” in Proceedings of the IEEE/CVF international conference on computer vision
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
K. An, Q. Chen, C. Deng, Z. Du, C. Gao, Z. Gao, Y. Gu, T. He, H. Hu, K. Hu, et al
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
K.-C. Wang, D. Ostashev, Y. Fang, S. Tulyakov, and K. Aberman, “Moa: Mixture-of-attention for subject-context disentanglement in personalized image generation,” in SIGGRAPH Asia 2024 Conference Papers
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Z. Su, Z. Lin, X. Bai, X. Wu, Y. Xiong, H. Lian, G. Ma, H. Chen, G. Ding, W. Zhou, et al
2024
Later among the works it cites.
E. Pedicir, L. Miller, and L. Robinson, “Novel token-level recurrent routing for enhanced mixture-of-experts performance,” Authorea Preprints
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
A. Wang, X. Sun, R. Xie, S. Li, J. Zhu, Z. Yang, P. Zhao, J. Han, Z. Kang, D. Wang, et al
2024
Later among the works it cites.
A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, et al
2024
Later among the works it cites.
J. Yao, Q. Anthony, A. Shafi, H. Subramoni, and D. K. D. Panda, “Exploiting inter-layer expert affinity for accelerating mixture-of-experts model inference,” in 2024 IEEE International Parallel and Distributed Processing Symposium (IPDPS)
2024
Later among the works it cites.
J. Yu, Y. Zhuge, L. Zhang, P. Hu, D. Wang, H. Lu, and Y. He, “Boosting continual learning of vision-language models via mixture-of-experts adapters,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2024
Later among the works it cites.
R. Zhang, A. Cheng, Y. Luo, G. Dai, H. Yang, J. Liu, R. Xu, L. Du, Y. Du, Y. Jiang, et al
2024
Later among the works it cites.
2024
Later among the works it cites.
X. Yu, T. Rosing, and Y. Guo, “Evolve: Enhancing unsupervised continual learning with multiple experts,” in 2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)
2024
Later among the works it cites.
L. Li, Z. Wu, and Y. Ji, “Mote: Mixture of task-specific experts for pre-trained model-based continual learning,” Available at SSRN 5035279
2024
Later among the works it cites.
2024
Later among the works it cites.
M. Le, H. Nguyen, T. Nguyen, T. Pham, L. Ngo, N. Ho, et al
2024
Later among the works it cites.
2024
Later among the works it cites.
Series Title: Lecture Notes in Computer Science
Q. Chen, L. Zhu, H. He, X. Zhang, S. Zeng, Q. Ren, and Y. Lu, “Low-rank mixture-of-experts for continual medical image segmentation,” in Medical Image Computing and Computer Assisted Intervention – MICCAI 2024 · 2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Yang, P.-T. Jiang, Q. Hou, H. Zhang, J. Chen, and B. Li, “Multi-task dense prediction via mixture of low-rank experts,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
W. Song, H. Zhao, P. Ding, C. Cui, S. Lyu, Y. Fan, and D. Wang, “Germ: A generalist robotic model with mixture-of-experts for quadruped robot,” in 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
K. Li, M. Cucuringu, L. Sánchez-Betancourt, and T. Willi, “Mixtures of experts for scaling up neural networks in order execution,” in Proceedings of the 5th ACM International Conference on AI in Finance
2024
Later among the works it cites.
V. Prasad, A. Kshirsagar, D. K. R. Stock-Homburg, J. Peters, and G. Chalvatzaki, “Moveint: Mixture of variational experts for learning human-robot interactions from demonstrations,” IEEE Robotics and Automation Letters
2024
Later among the works it cites.
H. Zeng, M. Xu, T. Zhou, X. Wu, J. Kang, Z. Cai, and D. Niyato, “One-shot-but-not-degraded federated learning,” in Proceedings of the 32nd ACM International Conference on Multimedia
2024
Later among the works it cites.
2024
Later among the works it cites.
H. Nguyen, T. Nguyen, K. Nguyen, and N. Ho, “Towards convergence rates for parameter estimation in gaussian-gated mixture of experts,” in International Conference on Artificial Intelligence and Statistics
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Jain, H. Behl, Z. Kira, and V. Vineet, “Damex: Dataset-aware mixture-of-experts for visual understanding of mixture-of-datasets,” Advances in Neural Information Processing Systems
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
X. Cheng, Z. Zhu, X. Zhuang, Z. Chen, Z. Huang, and Y. Zou, “MoE-SLU: Towards ASR-robust spoken language understanding via mixture-of-experts,” in Findings of the Association for Computational Linguistics ACL 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
X. Sun, Y. Chen, Y. Huang, R. Xie, J. Zhu, K. Zhang, S. Li, Z. Yang, J. Han, X. Shu, et al
2024
Later among the works it cites.
A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, et al
2024
Later among the works it cites.
T. Wei, B. Zhu, L. Zhao, C. Cheng, B. Li, W. Lü, P. Cheng, J. Zhang, X. Zhang, L. Zeng, et al
2024
Later among the works it cites.
N. Gupta and J. Yip, “Dbrx: Creating an llm from scratch using databricks,” in Databricks Data Intelligence Platform: Unlocking the GenAI Revolution
2024
Later among the works it cites.
T. Zhu, X. Qu, D. Dong, J. Ruan, J. Tong, C. He, and Y. Cheng, “Llama-moe: Building mixture-of-experts from llama with continual pre-training,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
S. Wu, J. Luo, X. Chen, L. Li, X. Zhao, T. Yu, C. Wang, Y. Wang, F. Wang, W. Qiao, et al
2024
Later among the works it cites.
D. Dai, C. Deng, C. Zhao, R. Xu, H. Gao, D. Chen, J. Li, W. Zeng, X. Yu, Y. Wu, et al
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
A. Vats, R. Raja, V. Jain, and A. Chadha, “The evolution of mixture of experts: A survey from basics to breakthroughs,” Preprints (August 2024)
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2025
Closest in time.
V.-T. Tran, L. H. Khiem, and V. Q. Pham, “Revisiting sparse mixture of experts for resource-adaptive federated fine-tuning foundation models,” in ICLR 2025 Workshop on Modularity for Collaborative, Decentralized, and Continual Deep Learning
2025
Closest in time.
T. C. Fung and S. C. Tseung, “Mixture of experts models for multilevel data: Modelling framework and approximation theory,” Neurocomputing
2025
Closest in time.
L. Rossi, V. Bernuzzi, T. Fontanini, M. Bertozzi, and A. Prati, “Swin2-mose: A new single image supersolution model for remote sensing,” IET Image Processing
2025
Closest in time.
B. Lenz, O. Lieber, A. Arazi, A. Bergman, A. Manevich, B. Peleg, B. Aviram, C. Almagor, C. Fridman, D. Padnos, et al
2025
Closest in time.
2025
Closest in time.
V. Dimitri, B. Regina, and M. Alfonz, “A survey on mixture of experts: Advancements, challenges, and future directions,” Authorea Preprints
2025
Closest in time.
Towards mixture of task-intensive experts for multi-task recommendation
X. Cai, Y. Lu, H. Lu, and Y. Ding · 2025
Closest in time.