Fetching the paper…

CLIP$^2$: Contrastive Language-Image-Point Pretraining from Real-World Point Cloud Data · Around