Fetching the paper…
Reading the bibliography…
We present a limited empirical study of scaling laws for transfer learning in transformer models.
Sanjeev Arora et al · 1901
Earlier work this paper cites.
Atsushi Nitanda, Geoffrey Chinot and Taiji Suzuki · 1905
Earlier work this paper cites.
“A Constructive Prediction of the Generalization Error Across Scales”, 2019
Jonathan. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov and Nir Shavit · 1909
Earlier work this paper cites.
“Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer”, 2023
Colin Raffel et al · 1910
Earlier work this paper cites.
“A Model of Inductive Bias Learning”
J. Baxter · 2000
Earlier work this paper cites.
“Scaling Laws for Neural Language Models”, 2020
Jared Kaplan et al · 2001
Earlier work this paper cites.
“A Neural Scaling Law from the Dimension of the Data Manifold”, 2020
Utkarsh Sharma and Jared Kaplan · 2004
Earlier work this paper cites.
“Language Models are Few-Shot Learners”, 2020
Tom. Brown et al · 2005
Earlier work this paper cites.
“On the Theory of Transfer Learning: The Importance of Task Diversity”, 2020
Nilesh Tripuraneni, Michael. Jordan and Chi Jin · 2006
Earlier work this paper cites.
“The Enron Email Dataset” Provided by Carnegie Mellon University’s School of Computer Science, Available for research use, 2009
William. Cohen · 2009
Earlier work this paper cites.
“Scaling Laws for Autoregressive Generative Modeling”, 2020
Tom Henighan et al · 2010
Earlier work this paper cites.
“The Benefit of Multitask Representation Learning”, 2016
Andreas Maurer, Massimiliano Pontil and Bernardino Romera-Paredes · 2016
Earlier work this paper cites.
“Deep Learning Scaling is Predictable, Empirically”, 2017
Joel Hestness et al · 2017
Cited alongside, same era.
“Domain Randomization for Sim2Real Transfer”
Lilian Weng · 2019
Cited alongside, same era.
“The Pile: An 800GB Dataset of Diverse Text for Language Modeling”, 2020
Leo Gao et al · 2020
Cited alongside, same era.
“Cat Assembly and Gene Annotation” Assembly: Felis_catus_9.0, INSDC Assembly GCA_000181335.4, Nov 2017. Genebuild last updated/patched September 2020. Database version 111.9. Ensembl release 111 - January 2024., 2020
Genome Sequencing Center at Washington University School of Medicine · 2020
Cited alongside, same era.
“Explaining Neural Scaling Laws”, 2021
Yasaman Bahri et al · 2021
“Scaling Laws for Generative Mixed-Modal Language Models”, 2023
Armen Aghajanyan et al · 2023
Later among the works it cites.
“Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling”, 2023
Stella Biderman et al · 2023
Later among the works it cites.
“Few-shot Fine-tuning vs. In-context Learning: A Fair Comparison and Evaluation”, 2023
Marius Mosbach et al · 2023
Later among the works it cites.
“Gemini: A Family of Highly Capable Multimodal Models”, 2023
Gemini Team et al · 2023
Later among the works it cites.
“Scaling Laws Literature Review” Accessed: 2023-9-12, 2023
Pablo Villalobos · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Scaling Laws for Transfer”, 2021
Danny Hernandez, Jared Kaplan, Tom Henighan and Sam McCandlish · 2021
Cited alongside, same era.
“Lora: Low-rank adaptation of large language models”
Edward Hu et al · 2021
Cited alongside, same era.
“A Scaling Law for Synthetic-to-Real Transfer: How Much Is Your Pre-training Effective?”, 2021
Hiroaki Mikami et al · 2021
Cited alongside, same era.
“When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method”, 2024
Biao Zhang, Zhongtao Liu, Colin Cherry and Orhan Firat · 2021
Cited alongside, same era.
“On the Opportunities and Risks of Foundation Models”, 2022
Rishi Bommasani et al · 2022
Cited alongside, same era.
“Training Compute-Optimal Large Language Models”, 2022
Jordan Hoffmann et al · 2022
Cited alongside, same era.
“Will we run out of data? An analysis of the limits of scaling datasets in Machine Learning”, 2022
Pablo Villalobos et al · 2022
Cited alongside, same era.
Later among the works it cites.
“Llama 3 Model Card”, 2024
AI@Meta · 2024
Closest in time.
“Chinchilla Scaling: A replication attempt”, 2024
Tamay Besiroglu, Ege Erdil, Matthew Barnett and Josh You · 2024
Closest in time.
“Parameter, Compute and Data Trends in Machine Learning” Accessed: 2024-05-22, 2024
Epoch AI · 2024
Closest in time.
“Statistics for Ecologists: A Frequentist and Bayesian Treatment of Modern Regression Models”
John. Fieberg · 2024
Closest in time.
Albert. Jiang et al · 2024
Closest in time.
“GPT-4 Technical Report”, 2024
OpenAI et al · 2024
Closest in time.
“Gemini: A Family of Highly Capable Multimodal Models”, 2024
Gemini Team et al · 2024
Closest in time.