Fetching the paper…

WIT: Wikipedia-based Image Text Dataset for Multimodal Multilingual Machine Learning · Around