Fetching the paper…

Learning to Generate Grounded Visual Captions without Localization Supervision · Around