Fetching the paper…

SwinBERT: End-to-End Transformers with Sparse Attention for Video Captioning · Around