2022

Instruction-Following Agents with Multimodal Transformer

Liu, Hao, Lee, Lisa, Lee, Kimin et al.

Understand

Humans are excellent at understanding language and vision to accomplish a wide range of tasks.

  • In contrast, creating general instruction-following embodied agents remains a difficult challenge.
  • Prior work that uses pure language-only models lack visual grounding, making it difficult to connect language instructions with visual observations.
  • On the other hand, methods that use pre-trained multimodal models typically come with divided language and visual representations, requiring designing specialized network architecture to fuse them together.

Reading the bibliography…