Fetching the paper…

Implicit Temporal Modeling with Learnable Alignment for Video Recognition · Around