Fetching the paper…

Multimodal Transformer with Multi-View Visual Representation for Image Captioning · Around