2020

Behind the Scene: Revealing the Secrets of Pre-trained Vision-and-Language Models

Cao, Jize, Gan, Zhe, Cheng, Yu et al.

Understand

Recent Transformer-based large-scale pre-trained models have revolutionized vision-and-language (V+L) research.

  • Models such as ViLBERT, LXMERT and UNITER have significantly lifted state of the art across a wide range of V+L benchmarks with joint image-text pre-training.
  • However, little is known about the inner mechanisms that destine their impressive success.
  • To reveal the secrets behind the scene of these powerful models, we present VALUE (Vision-And-Language Understanding Evaluation), a set of meticulously designed probing tasks (e.g., Visual Coreference Resolution, Visual Relation Detection, Linguistic Probing Tasks) generalizable to standard pre-trained V+L models, aiming to decipher the inner workings of multimodal pre-training (e.g., the implicit knowledge garnered in individual attention heads, the inherent cross-modal alignment learned through contextualized multimodal embeddings).

Reading the bibliography…