2023

$A^2$Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models

Chen, Peihao, Sun, Xinyu, Zhi, Hongyan et al.

Understand

We study the task of zero-shot vision-and-language navigation (ZS-VLN), a practical yet challenging problem in which an agent learns to navigate following a path described by language instructions without requiring any path-instruction annotation data.

  • Normally, the instructions have complex grammatical structures and often contain various action descriptions (e.g., "proceed beyond", "depart from").
  • How to correctly understand and execute these action demands is a critical problem, and the absence of annotated data makes it even more challenging.
  • Note that a well-educated human being can easily understand path instructions without the need for any special training.

Reading the bibliography…