Fetching the paper…

VELMA: Verbalization Embodiment of LLM Agents for Vision and Language Navigation in Street View · Around