2023

Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design

Alabdulmohsin, Ibrahim, Zhai, Xiaohua, Kolesnikov, Alexander et al.

Understand

Scaling laws have been recently employed to derive compute-optimal model size (number of parameters) for a given compute duration.

  • We advance and refine such methods to infer compute-optimal model shapes, such as width and depth, and successfully implement this in vision transformers.
  • Our shape-optimized vision transformer, SoViT, achieves results competitive with models that exceed twice its size, despite being pre-trained with an equivalent amount of compute.
  • For example, SoViT-400m/14 achieves 90.3% fine-tuning accuracy on ILSRCV2012, surpassing the much larger ViT-g/14 and approaching ViT-G/14 under identical settings, with also less than half the inference cost.

Reading the bibliography…