Fetching the paper…

AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding · Around