Fetching the paper…

VITA-VLA: Efficiently Teaching Vision-Language Models to Act via Action Expert Distillation · Around