2021

M6-10T: A Sharing-Delinking Paradigm for Efficient Multi-Trillion Parameter Pretraining

Lin, Junyang, Yang, An, Bai, Jinze et al.

Understand

Recent expeditious developments in deep learning algorithms, distributed training, and even hardware design for large models have enabled training extreme-scale models, say GPT-3 and Switch Transformer possessing hundreds of billions or even trillions of parameters.

  • However, under limited resources, extreme-scale model training that requires enormous amounts of computes and memory footprint suffers from frustratingly low efficiency in model convergence.
  • In this paper, we propose a simple training strategy called "Pseudo-to-Real" for high-memory-footprint-required large models.
  • Pseudo-to-Real is compatible with large models with architecture of sequential layers.

Reading the bibliography…