Fetching the paper…

Leveraging per Image-Token Consistency for Vision-Language Pre-training · Around