Fetching the paper…

4D-VLA: Spatiotemporal Vision-Language-Action Pretraining with Cross-Scene Calibration · Around