Fetching the paper…

PyramidCLIP: Hierarchical Feature Alignment for Vision-language Model Pretraining · Around