2024

Sparkle: Mastering Basic Spatial Capabilities in Vision Language Models Elicits Generalization to Spatial Reasoning

Tang, Yihong, Qu, Ao, Wang, Zhaokai et al.

Understand

Vision language models (VLMs) perform well on many tasks but often fail at spatial reasoning, which is essential for navigation and interaction with physical environments.

  • Many spatial reasoning tasks depend on fundamental two-dimensional (2D) skills, yet our evaluation shows that state-of-the-art VLMs give implausible or incorrect answers to composite spatial problems, including simple pathfinding tasks that humans solve effortlessly.
  • To address this, we enhance 2D spatial reasoning in VLMs by training them only on basic spatial capabilities.
  • We first disentangle 2D spatial reasoning into three core components: direction comprehension, distance estimation, and localization.

Reading the bibliography…