Fetching the paper…
Reading the bibliography…
Recent research has highlighted the importance of dataset size in scaling language models.
“Dropout: A Simple Way to Prevent Neural Networks from Overfitting”
Nitish Srivastava et al · 1958
Earlier work this paper cites.
“Imagenet large scale visual recognition challenge”
Olga Russakovsky et al · 2015
Earlier work this paper cites.
“Deep networks with stochastic depth”
Gao Huang et al · 2016
Earlier work this paper cites.
“SQuAD: 100,000+ Questions for Machine Comprehension of Text”
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev and Percy Liang · 2016
Earlier work this paper cites.
“Rethinking the inception architecture for computer vision”
Christian Szegedy et al · 2016
Earlier work this paper cites.
“Decoupled weight decay regularization”
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Mostafa Dehghani et al · 2018
Earlier work this paper cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2018
Earlier work this paper cites.
“Albert: A lite bert for self-supervised learning of language representations”
Zhenzhong Lan et al · 2019
Earlier work this paper cites.
“Megatron-lm: Training multi-billion parameter language models using model parallelism”
Mohammad Shoeybi et al · 2019
Earlier work this paper cites.
“An image is worth 16x16 words: Transformers for image recognition at scale”
Alexey Dosovitskiy et al · 2020
Earlier work this paper cites.
“Scaling laws for neural language models”
Jared Kaplan et al · 2020
Earlier work this paper cites.
“Gshard: Scaling giant models with conditional computation and automatic sharding”
Dmitry Lepikhin et al · 2020
Earlier work this paper cites.
“Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer”
Colin Raffel et al · 2020
Cited alongside, same era.
“mT5: A massively multilingual pre-trained text-to-text transformer”
Linting Xue et al · 2020
Cited alongside, same era.
“Scalable and efficient moe training for multitask multilingual models”
Young Kim et al · 2021
Cited alongside, same era.
“Tapex: Table pre-training via learning a neural sql executor”
Qian Liu et al · 2021
Cited alongside, same era.
“Cross-token Modeling with Conditional Computation”
Yuxuan Lou, Fuzhao Xue, Zangwei Zheng and Yang You · 2021
“Unifying language learning paradigms”
Yi Tay et al · 2022
Later among the works it cites.
“Galactica: A large language model for science”
Ross Taylor et al · 2022
Later among the works it cites.
“Will we run out of data? An analysis of the limits of scaling datasets in Machine Learning”
Pablo Villalobos et al · 2022
Later among the works it cites.
“Insights into Pre-training via Simpler Synthetic Tasks”
Yuhuai Wu, Felix Li and Percy Liang · 2022
Later among the works it cites.
“One Student Knows All Experts Know: From Sparse to Dense”
Fuzhao Xue et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Scaling language models: Methods, analysis & insights from training gopher”
Jack Rae et al · 2021
Cited alongside, same era.
“Scale efficiently: Insights from pre-training and fine-tuning transformers”
Yi Tay et al · 2021
Cited alongside, same era.
“Palm: Scaling language modeling with pathways”
Aakanksha Chowdhery et al · 2022
Cited alongside, same era.
“Scaling Laws and Interpretability of Learning from Repeated Data”
Danny Hernandez et al · 2022
Cited alongside, same era.
“Training compute-optimal large language models”
Jordan Hoffmann et al · 2022
Cited alongside, same era.
“Self-Prompting Large Language Models for Open-Domain QA”
Junlong Li, Zhuosheng Zhang and Hai Zhao · 2022
Cited alongside, same era.
“Training language models to follow instructions with human feedback”
Long Ouyang et al · 2022
Cited alongside, same era.
“Go wider instead of deeper”
Fuzhao Xue et al · 2022
Later among the works it cites.
“Byt5: Towards a token-free future with pre-trained byte-to-byte models”
Linting Xue et al · 2022
Later among the works it cites.
“St-moe: Designing stable and transferable sparse expert models”
Barret Zoph et al · 2022
Later among the works it cites.
Rohan Anil et al · 2023
Closest in time.
“Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling”, 2023
Stella Biderman et al · 2023
Closest in time.
OpenAI · 2023
Closest in time.
“Llama: Open and efficient foundation language models”
Hugo Touvron et al · 2023
Closest in time.