Fetching the paper…

Human-centric Spatio-Temporal Video Grounding With Visual Transformers · Around