Fetching the paper…

MMFT-BERT: Multimodal Fusion Transformer with BERT Encodings for Visual Question Answering · Around