Fetching the paper…

GRILL: Grounded Vision-language Pre-training via Aligning Text and Image Regions · Around