Fetching the paper…
Reading the bibliography…
We investigate the rate at which algorithms for pre-training language models have improved since the advent of deep learning.
“Backpropagation applied to handwritten zip code recognition”
Yann LeCun et al · 1989
Earlier work this paper cites.
“The Penn treebank: an overview”
Ann Taylor, Mitchell Marcus and Beatrice Santorini · 2003
Earlier work this paper cites.
“Language Models are Few-Shot Learners”
Tom. Brown et al · 2005
Earlier work this paper cites.
“A brief history of linear and mixed-integer programming computation”
Robert Bixby · 2012
Earlier work this paper cites.
“Pointer sentinel mixture models”
Stephen Merity, Caiming Xiong, James Bradbury and Richard Socher · 2016
Earlier work this paper cites.
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“Adaptive Input Representations for Neural Language Modeling”
Alexei Baevski and Michael Auli · 2018
Earlier work this paper cites.
“Direct Output Connection for a High-Rank Language Model”
Sho Takase, Jun Suzuki and Masaaki Nagata · 2018
Earlier work this paper cites.
“A survey on neural network language models”
Kun Jing and Jungang Xu · 2019
Earlier work this paper cites.
“Language Models are Unsupervised Multitask Learners”, 2019
Alec Radford et al · 2019
Earlier work this paper cites.
“Fast transformer decoding: One write-head is all you need”
Noam Shazeer · 2019
Earlier work this paper cites.
“Language models with transformers”
Chenguang Wang, Mu Li and Alexander Smola · 2019
Earlier work this paper cites.
“A time leap challenge for SAT-solving”
Johannes Fichte, Markus Hecher and Stefan Szeider · 2020
Earlier work this paper cites.
“The Pile: An 800GB dataset of diverse text for language modeling”
Leo Gao et al · 2020
Earlier work this paper cites.
“Measuring the algorithmic efficiency of neural networks”
Danny Hernandez and Tom Brown · 2020
Earlier work this paper cites.
“Scaling laws for neural language models”
Jared Kaplan et al · 2020
Earlier work this paper cites.
“Training Verifiers to Solve Math Word Problems”
Karl Cobbe et al · 2021
Earlier work this paper cites.
“Measuring progress in deep reinforcement learning sample efficiency”
Florian Dorner · 2021
Earlier work this paper cites.
“Scaling Laws for Acoustic Models”
Jasha Droppo and Oguz. Elibol · 2021
Cited alongside, same era.
“A framework for few-shot language model evaluation”
Leo Gao et al · 2021
Cited alongside, same era.
“Scaling Language Models: Methods, Analysis & Insights from Training Gopher”
Jack. Rae et al · 2021
Cited alongside, same era.
“How fast do algorithms improve?[point of view]”
Yash Sherry and Neil Thompson · 2021
Cited alongside, same era.
“Neuro-symbolic language modeling with automaton-augmented retrieval”
Uri Alon et al · 2022
Cited alongside, same era.
“FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness”
Tri Dao et al · 2022
Cited alongside, same era.
“AI capabilities can be significantly improved without expensive retraining”
Tom Davidson, Jean-Stanislas Denain, Pablo Villalobos and Guillem Bas · 2023
Later among the works it cites.
“Mamba: Linear-Time Sequence Modeling with Selective State Spaces”
Albert Gu and Tri Dao · 2023
Later among the works it cites.
Suriya Gunasekar et al · 2023
Later among the works it cites.
“Perplexity of fixed-length models” [Online; accessed 14-Nov-2023], https://huggingface.co/docs/transformers/perplexity , 2023
Hugging Face · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Algorithmic progress in computer vision”
Ege Erdil and Tamay Besiroglu · 2022
Cited alongside, same era.
“Training Compute-Optimal Large Language Models”
Jordan Hoffmann et al · 2022
Cited alongside, same era.
“Towards Reasoning in Large Language Models: A Survey”
Jie Huang and Kevin-Chuan Chang · 2022
Cited alongside, same era.
“Deep Neural Nets: 33 years ago and 33 years from now” [Online; accessed 21-July-2022], http://karpathy.github.io/2022/03/14/lecun1989/ , 2022
Andrej Karpathy · 2022
Cited alongside, same era.
“Progress in mathematical programming solvers from 2001 to 2020”
Thorsten Koch, Timo Berthold, Jaap Pedersen and Charlie Vanaret · 2022
Cited alongside, same era.
“Competition-level code generation with alphacode”
Yujia Li et al · 2022
Cited alongside, same era.
Jean Kaddour et al · 2023
Later among the works it cites.
“AlphaCode 2 Technical Report”, 2023
Rémi Leblond et al · 2023
Later among the works it cites.
“When less is more: Investigating data pruning for pretraining llms at scale”
Max Marion et al · 2023
Later among the works it cites.
“Augmented language models: a survey”
Grégoire Mialon et al · 2023
Later among the works it cites.
“Scaling Data-Constrained Language Models”
Niklas Muennighoff et al · 2023
Later among the works it cites.
“GPT-4 Technical Report”, 2023
OpenAI · 2023
Later among the works it cites.
“GPT-4 Architecture, Infrastructure, Training Dataset, Costs, Vision, MoE”, 2023
Dylan Patel and Gerald Wong · 2023
Later among the works it cites.
“In the long (context) run”, 2023
Harm de Vries · 2023
Later among the works it cites.
“Effective Long-Context Scaling of Foundation Models”
Wenhan Xiong et al · 2023
Later among the works it cites.
“A survey of large language models”
Wayne Zhao et al · 2023
Later among the works it cites.
“Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context”, 2024
Google Gemini · 2024
Closest in time.
“World Model on Million-Length Video And Language With RingAttention”
Hao Liu, Wilson Yan, Matei Zaharia and Pieter Abbeel · 2024
Closest in time.
“Solving olympiad geometry without human demonstrations”
Trieu Trinh et al · 2024
Closest in time.