Fetching the paper…
Reading the bibliography…
To help the open-source community have a better understanding of Mixture-of-Experts (MoE) based large language models (LLMs), we train and release OpenMoE, a series of fully open-sourced and reproducible decoder-only MoE LLMs, ranging from 650M to 34B parameters and trained on up to over 1T tokens.
“Findings of the 2016 Conference on Machine Translation”
Ondrej Bojar et al · 2016
Earlier work this paper cites.
“TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension”
Mandar Joshi, Eunsol Choi, Daniel Weld and Luke Zettlemoyer · 2017
Earlier work this paper cites.
“Outrageously large neural networks: The sparsely-gated mixture-of-experts layer”
Noam Shazeer et al · 2017
Earlier work this paper cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding”
Jacob-Wei Kenton and Lee Toutanova · 2019
Earlier work this paper cites.
“Roberta: A robustly optimized bert pretraining approach”
Yinhan Liu et al · 2019
Earlier work this paper cites.
“Megatron-lm: Training multi-billion parameter language models using model parallelism”
Mohammad Shoeybi et al · 2019
Earlier work this paper cites.
“Language models are few-shot learners”
Tom Brown et al · 2020
Earlier work this paper cites.
“An image is worth 16x16 words: Transformers for image recognition at scale”
Alexey Dosovitskiy et al · 2020
Earlier work this paper cites.
“Measuring massive multitask language understanding”
Dan Hendrycks et al · 2020
Earlier work this paper cites.
“Gshard: Scaling giant models with conditional computation and automatic sharding”
Dmitry Lepikhin et al · 2020
Earlier work this paper cites.
“Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer”
Colin Raffel et al · 2020
Earlier work this paper cites.
“Glu variants improve transformer”
Noam Shazeer · 2020
Earlier work this paper cites.
“CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data”
Guillaume Wenzek et al · 2020
Earlier work this paper cites.
“Efficient large scale language modeling with mixtures of experts”
Mikel Artetxe et al · 2021
Earlier work this paper cites.
“Evaluating Large Language Models Trained on Code”, 2021
Mark Chen et al · 2021
Earlier work this paper cites.
“Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity”
William Fedus, Barret Zoph and Noam Shazeer · 2021
Earlier work this paper cites.
“Base layers: Simplifying training of large, sparse models”
Mike Lewis et al · 2021
Earlier work this paper cites.
“Cross-token Modeling with Conditional Computation”
Yuxuan Lou, Fuzhao Xue, Zangwei Zheng and Yang You · 2021
Earlier work this paper cites.
“Scaling language models: Methods, analysis & insights from training gopher”
Jack Rae et al · 2021
Earlier work this paper cites.
“Scaling vision with sparse mixture of experts”
Carlos Riquelme et al · 2021
Earlier work this paper cites.
“Hash layers for large sparse models”
Stephen Roller, Sainbayar Sukhbaatar and Jason Weston · 2021
Earlier work this paper cites.
“GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model”, https://github.com/kingoflolz/mesh-transformer-jax , 2021
Ben Wang and Aran Komatsuzaki · 2021
Cited alongside, same era.
“GSPMD: general and scalable parallelization for ML computation graphs”
Yuanzhong Xu et al · 2021
Cited alongside, same era.
“Efficient training of language models to fill in the middle”
Mohammad Bavarian et al · 2022
Cited alongside, same era.
“Palm: Scaling language modeling with pathways”
Aakanksha Chowdhery et al · 2022
Cited alongside, same era.
“Glam: Efficient scaling of language models with mixture-of-experts”
Nan Du et al · 2022
Cited alongside, same era.
“RedPajama: An Open Source Recipe to Reproduce LLaMA training dataset”, 2023
Together Computer · 2023
Later among the works it cites.
“Megablocks: Efficient sparse training with mixture-of-experts”
Trevor Gale, Deepak Narayanan, Cliff Young and Matei Zaharia · 2023
Later among the works it cites.
“A framework for few-shot language model evaluation”
Leo Gao et al · 2023
Later among the works it cites.
“OpenLLaMA: An Open Reproduction of LLaMA”, 2023
Xinyang Geng and Hao Liu · 2023
Later among the works it cites.
“Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints”
Aran Komatsuzaki et al · 2023
Later among the works it cites.
“StarCoder: may the source be with you!”
Raymond Li et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“How does GPT Obtain its Ability? Tracing Emergent Abilities of Language Models to their Sources”
Hao Fu Yao; and Tushar Khot · 2022
Cited alongside, same era.
“Training compute-optimal large language models”
Jordan Hoffmann et al · 2022
Cited alongside, same era.
“The Stack: 3 TB of permissively licensed source code”
Denis Kocetkov et al · 2022
Cited alongside, same era.
“Self-Prompting Large Language Models for Open-Domain QA”
Junlong Li, Zhuosheng Zhang and Hai Zhao · 2022
Cited alongside, same era.
“Multimodal contrastive learning with limoe: the language-image mixture of experts”
Basil Mustafa et al · 2022
Cited alongside, same era.
“Ul2: Unifying language learning paradigms”
Yi Tay et al · 2022
Cited alongside, same era.
“Unifying language learning paradigms”
Yi Tay et al · 2022
Cited alongside, same era.
Later among the works it cites.
Erik Nijkamp et al · 2023
Later among the works it cites.
“From sparse to soft mixtures of experts”
Joan Puigcerver, Carlos Riquelme, Basil Mustafa and Neil Houlsby · 2023
Later among the works it cites.
“Code llama: Open foundation models for code”
Baptiste Roziere et al · 2023
Later among the works it cites.
“Dolma: An Open Corpus of Three Trillion Tokens for Language Model Pretraining Research”
Luca Soldaini et al · 2023
Later among the works it cites.
“LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training”, 2023
LLaMA-MoE Team · 2023
Later among the works it cites.
“Llama: Open and efficient foundation language models”
Hugo Touvron et al · 2023
Later among the works it cites.
“Deep long-tailed learning: A survey”
Yifan Zhang et al · 2023
Later among the works it cites.
“Judging LLM-as-a-judge with MT-Bench and Chatbot Arena”
Lianmin Zheng et al · 2023
Later among the works it cites.
“Brainformers: Trading simplicity for efficiency”
Yanqi Zhou et al · 2023
Later among the works it cites.
“(InThe)WildChat: 570K ChatGPT Interaction Logs In The Wild”
Anonymous · 2024
Closest in time.
“DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models”
Damai Dai et al · 2024
Closest in time.
Albert Jiang et al · 2024
Closest in time.
“Roformer: Enhanced transformer with rotary position embedding”
Jianlin Su et al · 2024
Closest in time.
“TinyLlama: An Open-Source Small Language Model”, 2024
Peiyuan Zhang, Guangtao Zeng, Tianduo Wang and Wei Lu · 2024
Closest in time.