2021

ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision

Kim, Wonjae, Son, Bokyung, Kim, Ildoo

Understand

Vision-and-Language Pre-training (VLP) has improved performance on various joint vision-and-language downstream tasks.

  • Current approaches to VLP heavily rely on image feature extraction processes, most of which involve region supervision (e.g., object detection) and the convolutional architecture (e.g., ResNet).
  • Although disregarded in the literature, we find it problematic in terms of both (1) efficiency/speed, that simply extracting input features requires much more computation than the multimodal interaction steps; and (2) expressive power, as it is upper bounded to the expressive power of the visual embedder and its predefined visual vocabulary.
  • In this paper, we present a minimal VLP model, Vision-and-Language Transformer (ViLT), monolithic in the sense that the processing of visual inputs is drastically simplified to just the same convolution-free manner that we process textual inputs.

Reading the bibliography…