Fetching the paper…
Reading the bibliography…
While GPUs are responsible for training the vast majority of state-of-the-art deep learning models, the implications of their architecture are often overlooked when designing new deep learning (DL) models.
“Language Models are Few-Shot Learners”
Tom Brown et al · 1901
Earlier work this paper cites.
Yuhsiang Tsai, Terry Cojean and Hartwig Anzt · 2008
Earlier work this paper cites.
“Attention is All You Need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“Harnessing GPU Tensor Cores for Fast FP16 Arithmetic to Speed up Mixed-Precision Iterative Refinement Solvers”
Azzam Haidar, Stanimire Tomov, Jack Dongarra and Nicholas. Higham · 2018
Earlier work this paper cites.
“Modeling Deep Learning Accelerator Enabled GPUs”
Md Raihan, Negar Goli and Tor. Aamodt · 2018
Earlier work this paper cites.
“BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2019
Earlier work this paper cites.
“Roberta: A robustly optimized bert pretraining approach”
Yinhan Liu et al · 2019
Earlier work this paper cites.
“A survey of techniques for optimizing deep learning on GPUs”
Sparsh Mittal and Shraiysh Vaishay · 2019
Earlier work this paper cites.
“Language models are unsupervised multitask learners”
Alec Radford et al · 2019
Earlier work this paper cites.
“Megatron-LM: Training Multi-Billion Parameter Language Models using GPU Model Parallelism”
Mohammad Shoeybi et al · 2019
Earlier work this paper cites.
“Benchmarking TPU, GPU, and CPU Platforms for Deep Learning”
Yu Wang, Gu-Yeon Wei and David. Brooks · 2019
Earlier work this paper cites.
“XSP: Across-Stack Profiling and Analysis of Machine Learning Models on GPUs”
C. Li et al · 2020
Earlier work this paper cites.
“Exploring the limits of transfer learning with a unified text-to-text transformer”
Colin Raffel et al · 2020
Earlier work this paper cites.
“Glu variants improve transformer”
Noam Shazeer · 2020
Earlier work this paper cites.
“Demystifying Tensor Cores to Optimize Half-Precision Matrix Multiply”
Da Yan, Wei Wang and Xiaowen Chu · 2020
Earlier work this paper cites.
“GPT-Neo: Large scale autoregressive language modeling with mesh-tensorflow”
Sid Black et al · 2021
Cited alongside, same era.
“TurboTransformers: An Efficient GPU Serving System for Transformer Models”
Jiarui Fang, Yang Yu, Chengduo Zhao and Jie Zhou · 2021
Cited alongside, same era.
“FasterTransformer”, https://github.com/NVIDIA/FasterTransformer , 2021
2021
Cited alongside, same era.
Zachary Nado et al · 2021
Cited alongside, same era.
“Efficient Large-Scale Language Model Training on GPU Clusters using Megatron-LM”
Deepak Narayanan et al · 2021
Cited alongside, same era.
“Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation”
Shaden Smith et al · 2022
Later among the works it cites.
“Releasing 3B and 7B RedPajama-INCITE family of models including base, instruction-tuned & chat models”, 2023
Together AI · 2023
Later among the works it cites.
“GPT-NeoX: Large Scale Autoregressive Language Modeling in PyTorch”, GitHub Repo, 2023
Alex Andonian et al · 2023
Later among the works it cites.
“Pythia: A suite for analyzing large language models across training and scaling”
Stella Biderman et al · 2023
Later among the works it cites.
“RedPajama: an Open Dataset for Training Large Language Models”, 2023
Together Computer · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ofir Press, Noah Smith and Mike Lewis · 2021
Cited alongside, same era.
“Roformer: Enhanced transformer with rotary position embedding”
Jianlin Su et al · 2021
Cited alongside, same era.
“GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model”, 2021
Ben Wang and Aran Komatsuzaki · 2021
Cited alongside, same era.
“Comparative evaluation of deep learning workloads for leadership-class systems”
Junqi Yin et al · 2021
Cited alongside, same era.
Reza Aminabadi et al · 2022
Cited alongside, same era.
“PaLM: Scaling Language Modeling with Pathways” Version 5
Aakanksha Chowdhery et al · 2022
Cited alongside, same era.
“Flashattention: Fast and memory-efficient exact attention with io-awareness”
Tri Dao et al · 2022
Cited alongside, same era.
Tri Dao · 2023
Later among the works it cites.
Nolan Dey et al · 2023
Later among the works it cites.
“Let’s talk about a detail that occurs during PyTorch 2.0’s codegen - tiling.”
Horace He · 2023
Later among the works it cites.
“The most dramatic optimization to nanoGPT so far ( 25% speedup) is to simply increase vocab size from 50257 to 50304 (nearest multiple of 64).”
Andrej Karpathy · 2023
Later among the works it cites.
“MLPerf” Accessed: August 24, 2026, https://mlperf.org/, 2023
2023
Later among the works it cites.
“Matrix Multiplication Background”, User’s Guide — NVIDIA Docs, 2023
NVIDIA · 2023
Later among the works it cites.
“OLCF6 Technical Requirements and Benchmarks”, 2023
OLCF · 2023
Later among the works it cites.
“ByteTransformer: A High-Performance Transformer Boosted for Variable-Length Inputs”
Y. Zhai et al · 2023
Later among the works it cites.
“Opt: Open pre-trained transformer language models”
Susan Zhang et al · 2048
Closest in time.