Fetching the paper…

LXMERT: Learning Cross-Modality Encoder Representations from Transformers · Around