Fetching the paper…

Coarse-to-Fine Vision-Language Pre-training with Fusion in the Backbone · Around