Fetching the paper…

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos · Around