2022

Generalized Decoding for Pixel, Image, and Language

Zou, Xueyan, Dou, Zi-Yi, Yang, Jianwei et al.

Understand

We present X-Decoder, a generalized decoding model that can predict pixel-level segmentation and language tokens seamlessly.

  • X-Decodert takes as input two types of queries: (i) generic non-semantic queries and (ii) semantic queries induced from text inputs, to decode different pixel-level and token-level outputs in the same semantic space.
  • With such a novel design, X-Decoder is the first work that provides a unified way to support all types of image segmentation and a variety of vision-language (VL) tasks.
  • Further, our design enables seamless interactions across tasks at different granularities and brings mutual benefits by learning a common and rich pixel-level visual-semantic understanding space, without any pseudo-labeling.

Reading the bibliography…