2018

Adversarial Semantic Alignment for Improved Image Captions

Dognin, Pierre L., Melnyk, Igor, Mroueh, Youssef et al.

Understand

In this paper we study image captioning as a conditional GAN training, proposing both a context-aware LSTM captioner and co-attentive discriminator, which enforces semantic alignment between images and captions.

  • We empirically focus on the viability of two training methods: Self-critical Sequence Training (SCST) and Gumbel Straight-Through (ST) and demonstrate that SCST shows more stable gradient behavior and improved results over Gumbel ST, even without accessing discriminator gradients directly.
  • We also address the problem of automatic evaluation for captioning models and introduce a new semantic score, and show its correlation to human judgement.
  • As an evaluation paradigm, we argue that an important criterion for a captioner is the ability to generalize to compositions of objects that do not usually co-occur together.

Reading the bibliography…