2020

Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Huang, Zhicheng, Zeng, Zhaoyang, Liu, Bei et al.

Understand

We propose Pixel-BERT to align image pixels with text by deep multi-modal transformers that jointly learn visual and language embedding in a unified end-to-end framework.

  • We aim to build a more accurate and thorough connection between image pixels and language semantics directly from image and sentence pairs instead of using region-based image features as the most recent vision and language tasks.
  • Our Pixel-BERT which aligns semantic connection in pixel and text level solves the limitation of task-specific visual representation for vision and language tasks.
  • It also relieves the cost of bounding box annotations and overcomes the unbalance between semantic labels in visual task and language semantic.

Reading the bibliography…