Fetching the paper…

PEVL: Position-enhanced Pre-training and Prompt Tuning for Vision-language Models · Around