Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have garnered unprecedented advancements across diverse fields, ranging from natural language processing to computer vision and beyond.
J. Ma, Z. Zhao, X. Yi, J. Chen, L. Hong, and E. H. Chi, “Modeling task relationships in multi-task learning with multi-gate mixture-of-experts,” in Proceedings of the 24th ACM SIGKDD international conference on knowledge discovery & data mining , 2018, pp. 1930–1939
1939
Earlier work this paper cites.
R. A. Jacobs, M. I. Jordan, S. J. Nowlan, and G. E. Hinton, “Adaptive mixtures of local experts,” Neural computation , vol. 3, no. 1, pp. 79–87, 1991
1991
Earlier work this paper cites.
M. I. Jordan and R. A. Jacobs, “Hierarchical mixtures of experts and the EM algorithm,” Neural computation , vol. 6, no. 2, pp. 181–214, 1994
1994
Earlier work this paper cites.
R. Collobert, S. Bengio, and Y. Bengio, “A parallel mixture of SVMs for very large scale problems,” Advances in Neural Information Processing Systems , vol. 14, 2001
2001
Earlier work this paper cites.
C. Rasmussen and Z. Ghahramani, “Infinite mixtures of Gaussian process experts,” Advances in neural information processing systems , vol. 14, 2001
2001
Earlier work this paper cites.
B. Shahbaba and R. Neal, “Nonlinear models using Dirichlet process mixtures.” Journal of Machine Learning Research , vol. 10, no. 8, 2009
2009
Earlier work this paper cites.
X. Glorot, A. Bordes, and Y. Bengio, “Deep sparse rectifier neural networks,” in Proceedings of the fourteenth international conference on artificial intelligence and statistics . JMLR Workshop and Conference Proceedings, 2011, pp. 315–323
2011
Earlier work this paper cites.
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng, “Multimodal deep learning,” in Proceedings of the 28th international conference on machine learning (ICML-11) , 2011, pp. 689–696
2011
Earlier work this paper cites.
S. E. Yuksel, J. N. Wilson, and P. D. Gader, “Twenty years of mixture of experts,” IEEE transactions on neural networks and learning systems , vol. 23, no. 8, pp. 1177–1193, 2012
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
L. Theis and M. Bethge, “Generative image modeling using spatial lstms,” Advances in neural information processing systems , vol. 28, 2015
2015
Earlier work this paper cites.
M. Deisenroth and J. W. Ng, “Distributed gaussian processes,” in International conference on machine learning . PMLR, 2015, pp. 1481–1490
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Almahairi, N. Ballas, T. Cooijmans, Y. Zheng, H. Larochelle, and A. Courville, “Dynamic capacity networks,” in International Conference on Machine Learning . PMLR, 2016, pp. 2549–2558
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
R. Aljundi, P. Chakravarty, and T. Tuytelaars, “Expert gate: Lifelong learning with a network of experts,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 3366–3375
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
S. Gross, M. Ranzato, and A. Szlam, “Hard mixtures of experts for large scale weakly supervised vision,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 6865–6873
2017
Earlier work this paper cites.
J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska et al. , “Overcoming catastrophic forgetting in neural networks,” Proceedings of the national academy of sciences , vol. 114, no. 13, pp. 3521–3526, 2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
N. Shazeer, Y. Cheng, N. Parmar, D. Tran, A. Vaswani, P. Koanantakool, P. Hawkins, H. Lee, M. Hong, C. Young et al. , “Mesh-tensorflow: Deep learning for supercomputers,” Advances in neural information processing systems , vol. 31, 2018
2018
Earlier work this paper cites.
T. Baltrušaitis, C. Ahuja, and L.-P. Morency, “Multimodal machine learning: A survey and taxonomy,” IEEE transactions on pattern analysis and machine intelligence , vol. 41, no. 2, pp. 423–443, 2018
2018
Earlier work this paper cites.
J. Lu, D. Batra, D. Parikh, and S. Lee, “Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
G. Lample, A. Sablayrolles, M. Ranzato, L. Denoyer, and H. Jégou, “Large memory layers with product keys,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Earlier work this paper cites.
B. Zhang and R. Sennrich, “Root mean square layer normalization,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Earlier work this paper cites.
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly, “Parameter-efficient transfer learning for NLP,” in International Conference on Machine Learning . PMLR, 2019, pp. 2790–2799
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Y. Huang, Y. Cheng, A. Bapna, O. Firat, D. Chen, M. Chen, H. Lee, J. Ngiam, Q. V. Le, Y. Wu et al. , “Gpipe: Efficient training of giant neural networks using pipeline parallelism,” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
D. Narayanan, A. Harlap, A. Phanishayee, V. Seshadri, N. R. Devanur, G. R. Ganger, P. B. Gibbons, and M. Zaharia, “PipeDream: generalized pipeline parallelism for DNN training,” in Proceedings of the 27th ACM symposium on operating systems principles , 2019, pp. 1–15
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
X. Wang, F. Yu, L. Dunlap, Y.-A. Ma, R. Wang, A. Mirhoseini, T. Darrell, and J. E. Gonzalez, “Deep mixture of experts via shallow embedding,” in Uncertainty in artificial intelligence . PMLR, 2020, pp. 552–562
2020
Earlier work this paper cites.
N. Shazeer, “Glu variants improve transformer,” arXiv preprint arXiv:2002.05202 , 2020
2020
Earlier work this paper cites.
H. Tang, J. Liu, M. Zhao, and X. Gong, “Progressive layered extraction (ple): A novel multi-task learning (mtl) model for personalized recommendations,” in Proceedings of the 14th ACM Conference on Recommender Systems , 2020, pp. 269–278
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of machine learning research , vol. 21, no. 140, pp. 1–67, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He, “Zero: Memory optimizations toward training trillion parameter models,” in SC20: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 2020, pp. 1–16
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
L. Zhou, H. Palangi, L. Zhang, H. Hu, J. Corso, and J. Gao, “Unified vision-language pre-training for image captioning and vqa,” in Proceedings of the AAAI conference on artificial intelligence , vol. 34, no. 07, 2020, pp. 13 041–13 049
2020
Earlier work this paper cites.
C. Riquelme, J. Puigcerver, B. Mustafa, M. Neumann, R. Jenatton, A. Susano Pinto, D. Keysers, and N. Houlsby, “Scaling vision with sparse mixture of experts,” Advances in Neural Information Processing Systems , vol. 34, pp. 8583–8595, 2021
2021
Earlier work this paper cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 10 012–10 022
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
M. Lewis, S. Bhosale, T. Dettmers, N. Goyal, and L. Zettlemoyer, “Base layers: Simplifying training of large, sparse models,” in International Conference on Machine Learning . PMLR, 2021, pp. 6265–6274
2021
Earlier work this paper cites.
H. Hazimeh, Z. Zhao, A. Chowdhery, M. Sathiamoorthy, Y. Chen, R. Mazumder, L. Hong, and E. Chi, “Dselect-k: Differentiable selection in the mixture of experts with applications to multi-task learning,” Advances in Neural Information Processing Systems , vol. 34, pp. 29 335–29 347, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
S. Kudugunta, Y. Huang, A. Bapna, M. Krikun, D. Lepikhin, M.-T. Luong, and O. Firat, “Beyond Distillation: Task-level Mixture-of-Experts for Efficient Inference,” in Findings of the Association for Computational Linguistics: EMNLP 2021 , 2021, pp. 3577–3599
2021
Earlier work this paper cites.
S. Roller, S. Sukhbaatar, J. Weston et al. , “Hash layers for large sparse models,” Advances in Neural Information Processing Systems , vol. 34, pp. 17 555–17 566, 2021
2021
Earlier work this paper cites.
S. Zuo, X. Liu, J. Jiao, Y. J. Kim, H. Hassan, R. Zhang, J. Gao, and T. Zhao, “Taming Sparsely Activated Transformer with Stochastic Experts,” in International Conference on Learning Representations , 2021
2021
Earlier work this paper cites.
A. Fan, S. Bhosale, H. Schwenk, Z. Ma, A. El-Kishky, S. Goyal, M. Baines, O. Celebi, G. Wenzek, V. Chaudhary et al. , “Beyond english-centric multilingual machine translation,” Journal of Machine Learning Research , vol. 22, no. 107, pp. 1–48, 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
E. J. Hu, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen et al. , “LoRA: Low-Rank Adaptation of Large Language Models,” in International Conference on Learning Representations , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
X. L. Li and P. Liang, “Prefix-Tuning: Optimizing Continuous Prompts for Generation,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) , 2021, pp. 4582–4597
2021
Earlier work this paper cites.
J. Ren, S. Rajbhandari, R. Y. Aminabadi, O. Ruwase, S. Yang, M. Zhang, D. Li, and Y. He, “ { \{ Zero-offload } \} : Democratizing { \{ billion-scale } \} model training,” in 2021 USENIX Annual Technical Conference (USENIX ATC 21) , 2021, pp. 551–564
2021
Earlier work this paper cites.
S. Rajbhandari, O. Ruwase, J. Rasley, S. Smith, and Y. He, “Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning,” in Proceedings of the international conference for high performance computing, networking, storage and analysis , 2021, pp. 1–14
2021
Earlier work this paper cites.
D. Narayanan, M. Shoeybi, J. Casper, P. LeGresley, M. Patwary, V. Korthikanti, D. Vainbrand, P. Kashinkunti, J. Bernauer, B. Catanzaro et al. , “Efficient large-scale language model training on gpu clusters using megatron-lm,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , 2021, pp. 1–15
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
K. Zhou, J. Yang, C. C. Loy, and Z. Liu, “Learning to prompt for vision-language models,” International Journal of Computer Vision , vol. 130, no. 9, pp. 2337–2348, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
N. Du, Y. Huang, A. M. Dai, S. Tong, D. Lepikhin, Y. Xu, M. Krikun, Y. Zhou, A. W. Yu, O. Firat et al. , “Glam: Efficient scaling of language models with mixture-of-experts,” in International Conference on Machine Learning . PMLR, 2022, pp. 5547–5569
2022
Earlier work this paper cites.
W. Fedus, B. Zoph, and N. Shazeer, “Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,” Journal of Machine Learning Research , vol. 23, no. 120, pp. 1–39, 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
Y. Wang, S. Agarwal, S. Mukherjee, X. Liu, J. Gao, A. H. Awadallah, and J. Gao, “AdaMix: Mixture-of-Adaptations for Parameter-efficient Model Tuning,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing , Y. Goldberg, Z. Kozareva, and Y. Zhang, Eds. Abu Dhabi, United Arab Emirates: Association for Computational Linguistics, Dec. 2022, pp. 5744–5760. [Online]. Available: https://aclanthology.org/2022.emnlp-main.388
2022
Cited alongside, same era.
M. Zhai, J. He, Z. Ma, Z. Zong, R. Zhang, and J. Zhai, “ { \{ SmartMoE } \} : Efficiently Training { \{ Sparsely-Activated } \} Models through Combining Offline and Online Parallelization,” in 2023 USENIX Annual Technical Conference (USENIX ATC 23) , 2023, pp. 961–975
2023
Later among the works it cites.
T. Gale, D. Narayanan, C. Young, and M. Zaharia, “Megablocks: Efficient sparse training with mixture-of-experts,” Proceedings of Machine Learning and Systems , vol. 5, 2023
2023
Later among the works it cites.
N. Zheng, H. Jiang, Q. Zhang, Z. Han, L. Ma, Y. Yang, F. Yang, C. Zhang, L. Qiu, M. Yang et al. , “Pit: Optimization of dynamic sparse deep learning models via permutation invariant transformation,” in Proceedings of the 29th Symposium on Operating Systems Principles , 2023, pp. 331–347
2023
Later among the works it cites.
S. Shi, X. Pan, X. Chu, and B. Li, “Pipemoe: Accelerating mixture-of-experts through adaptive pipelining,” in IEEE INFOCOM 2023-IEEE Conference on Computer Communications . IEEE, 2023, pp. 1–10
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Komatsuzaki, J. Puigcerver, J. Lee-Thorp, C. R. Ruiz, B. Mustafa, J. Ainslie, Y. Tay, M. Dehghani, and N. Houlsby, “Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints,” in The Eleventh International Conference on Learning Representations , 2022
2022
Cited alongside, same era.
Z. Zhang, Y. Lin, Z. Liu, P. Li, M. Sun, and J. Zhou, “MoEfication: Transformer Feed-forward Layers are Mixtures of Experts,” in Findings of the Association for Computational Linguistics: ACL 2022 , 2022, pp. 877–890
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
F. Xue, Z. Shi, F. Wei, Y. Lou, Y. Liu, and Y. You, “Go wider instead of deeper,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 8, 2022, pp. 8779–8787
2022
Cited alongside, same era.
A. Clark, D. de Las Casas, A. Guy, A. Mensch, M. Paganini, J. Hoffmann, B. Damoc, B. Hechtman, T. Cai, S. Borgeaud et al. , “Unified scaling laws for routed language models,” in International conference on machine learning . PMLR, 2022, pp. 4057–4086
2022
Cited alongside, same era.
S. Rajbhandari, C. Li, Z. Yao, M. Zhang, R. Y. Aminabadi, A. A. Awan, J. Rasley, and Y. He, “Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale,” in International conference on machine learning . PMLR, 2022, pp. 18 332–18 346
2022
Cited alongside, same era.
D. Dai, L. Dong, S. Ma, B. Zheng, Z. Sui, B. Chang, and F. Wei, “StableMoE: Stable Routing Strategy for Mixture of Experts,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2022, pp. 7085–7095
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Jiang, P. Zhang, Y. Luo, C. Li, J. B. Kim, K. Zhang, S. Wang, X. Xie, and S. Kim, “AdaMCT: adaptive mixture of CNN-transformer for sequential recommendation,” in Proceedings of the 32nd ACM International Conference on Information and Knowledge Management , 2023, pp. 976–986
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Li, B. Hui, Z. Yin, M. Yang, F. Huang, and Y. Li, “Pace: Unified multi-modal dialogue pre-training with progressive and compositional experts,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2023, pp. 13 402–13 416
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
J. Ainslie, J. Lee-Thorp, M. de Jong, Y. Zemlyanskiy, F. Lebron, and S. Sanghai, “GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , 2023, pp. 4895–4901
2023
Later among the works it cites.
2023
Later among the works it cites.
O. Ostapenko, L. Caccia, Z. Su, N. Le Roux, L. Charlin, and A. Sordoni, “A Case Study of Instruction Tuning with Mixture of Parameter-Efficient Experts,” in NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following , 2023
2023
Later among the works it cites.
T. Chen, Z. Zhang, A. K. JAISWAL, S. Liu, and Z. Wang, “Sparse MoE as the New Dropout: Scaling Dense and Self-Slimmable Transformers,” in The Eleventh International Conference on Learning Representations , 2023. [Online]. Available: https://openreview.net/forum?id=w1hwFUb_81
2023
Later among the works it cites.
2023
Later among the works it cites.
P. Qi, X. Wan, G. Huang, and M. Lin, “Zero Bubble Pipeline Parallelism,” in The Twelfth International Conference on Learning Representations , 2023
2023
Later among the works it cites.
V. A. Korthikanti, J. Casper, S. Lym, L. McAfee, M. Andersch, M. Shoeybi, and B. Catanzaro, “Reducing activation recomputation in large transformer models,” Proceedings of Machine Learning and Systems , vol. 5, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with pagedattention,” in Proceedings of the 29th Symposium on Operating Systems Principles , 2023, pp. 611–626
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
xAI, “Grok-1,” March 2024. [Online]. Available: https://github.com/xai-org/grok-1
2024
Closest in time.
Databricks, “Introducing DBRX: A New State-of-the-Art Open LLM,” March 2024. [Online]. Available: https://www.databricks.com/blog/introducing-dbrx-new-state-art-open-llm
2024
Closest in time.
S. A. R. Team, “Snowflake Arctic: The Best LLM for Enterprise AI — Efficiently Intelligent, Truly Open,” April 2024. [Online]. Available: https://www.snowflake.com/blog/arctic-open-efficient-foundation-language-models-snowflake/
2024
Closest in time.
DeepSeek-AI, “DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model,” 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
X. Wu, S. Huang, and F. Wei, “Mixture of LoRA Experts,” in The Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https://openreview.net/forum?id=uWvKBCYh4S
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Q. Team, “Qwen1.5-MoE: Matching 7B Model Performance with 1/3 Activated Parameters”,” February 2024. [Online]. Available: https://qwenlm.github.io/blog/qwen-moe/
2024
Closest in time.
X. O. He, “Mixture of a million experts,” arXiv preprint arXiv:2407.04153 , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Qiu, Z. Huang, and J. Fu, “Unlocking emergent modularity in large language models,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , 2024, pp. 2638–2660
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Zhang, Y. Xia, H. Wang, D. Yang, C. Hu, X. Zhou, and D. Cheng, “MPMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism,” IEEE Transactions on Parallel and Distributed Systems , 2024
2024
Closest in time.
2024
Closest in time.
S. Shi, X. Pan, Q. Wang, C. Liu, X. Ren, Z. Hu, Y. Yang, B. Li, and X. Chu, “Schemoe: An extensible mixture-of-experts distributed training system with tasks scheduling,” in Proceedings of the Nineteenth European Conference on Computer Systems , 2024, pp. 236–249
2024
Closest in time.
2024
Closest in time.
Z. Zhang, S. Liu, J. Yu, Q. Cai, X. Zhao, C. Zhang, Z. Liu, Q. Liu, H. Zhao, L. Hu et al. , “M3oe: Multi-domain multi-task mixture-of experts recommendation framework,” in Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval , 2024, pp. 893–902
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Q. Team, “Introducing Qwen1.5,” February 2024. [Online]. Available: https://qwenlm.github.io/blog/qwen1.5/
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu, “Roformer: Enhanced transformer with rotary position embedding,” Neurocomputing , vol. 568, p. 127063, 2024
2024
Closest in time.