Fetching the paper…

Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks · Around