Fetching the paper…

StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling · Around