2021

WebQA: Multihop and Multimodal QA

Chang, Yingshan, Narang, Mridu, Suzuki, Hisami et al.

Understand

Scaling Visual Question Answering (VQA) to the open-domain and multi-hop nature of web searches, requires fundamental advances in visual representation learning, knowledge aggregation, and language generation.

  • In this work, we introduce WebQA, a challenging new benchmark that proves difficult for large-scale state-of-the-art models which lack language groundable visual representations for novel objects and the ability to reason, yet trivial for humans.
  • WebQA mirrors the way humans use the web: 1) Ask a question, 2) Choose sources to aggregate, and 3) Produce a fluent language response.
  • This is the behavior we should be expecting from IoT devices and digital assistants.

Reading the bibliography…