2021

Scaling Up Vision-Language Pre-training for Image Captioning

Hu, Xiaowei, Gan, Zhe, Wang, Jianfeng et al.

Understand

In recent years, we have witnessed significant performance boost in the image captioning task based on vision-language pre-training (VLP).

  • Scale is believed to be an important factor for this advance.
  • However, most existing work only focuses on pre-training transformers with moderate sizes (e.g., 12 or 24 layers) on roughly 4 million images.
  • In this paper, we present LEMON, a LargE-scale iMage captiONer, and provide the first empirical study on the scaling behavior of VLP for image captioning.

Reading the bibliography…