2020

Deconfounded Image Captioning: A Causal Retrospect

Yang, Xu, Zhang, Hanwang, Cai, Jianfei

Understand

Dataset bias in vision-language tasks is becoming one of the main problems which hinders the progress of our community.

  • Existing solutions lack a principled analysis about why modern image captioners easily collapse into dataset bias.
  • In this paper, we present a novel perspective: Deconfounded Image Captioning (DIC), to find out the answer of this question, then retrospect modern neural image captioners, and finally propose a DIC framework: DICv1.0 to alleviate the negative effects brought by dataset bias.
  • DIC is based on causal inference, whose two principles: the backdoor and front-door adjustments, help us review previous studies and design new effective models.

Reading the bibliography…