Fetching the paper…

VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation · Around