2020

Words aren't enough, their order matters: On the Robustness of Grounding Visual Referring Expressions

Akula, Arjun R, Gella, Spandana, Al-Onaizan, Yaser et al.

Understand

Visual referring expression recognition is a challenging task that requires natural language understanding in the context of an image.

  • We critically examine RefCOCOg, a standard benchmark for this task, using a human study and show that 83.7% of test instances do not require reasoning on linguistic structure, i.e., words are enough to identify the target object, the word order doesn't matter.
  • To measure the true progress of existing models, we split the test set into two sets, one which requires reasoning on linguistic structure and the other which doesn't.
  • Additionally, we create an out-of-distribution dataset Ref-Adv by asking crowdworkers to perturb in-domain examples such that the target object changes.

Reading the bibliography…