Fetching the paper…

A Better Use of Audio-Visual Cues: Dense Video Captioning with Bi-modal Transformer · Around