2021

Comprehensive Visual Question Answering on Point Clouds through Compositional Scene Manipulation

Yan, Xu, Yuan, Zhihao, Du, Yuhao et al.

Understand

Visual Question Answering on 3D Point Cloud (VQA-3D) is an emerging yet challenging field that aims at answering various types of textual questions given an entire point cloud scene.

  • To tackle this problem, we propose the CLEVR3D, a large-scale VQA-3D dataset consisting of 171K questions from 8,771 3D scenes.
  • Specifically, we develop a question engine leveraging 3D scene graph structures to generate diverse reasoning questions, covering the questions of objects' attributes (i.e., size, color, and material) and their spatial relationships.
  • Through such a manner, we initially generated 44K questions from 1,333 real-world scenes.

Reading the bibliography…