Fetching the paper…

Co-training Transformer with Videos and Images Improves Action Recognition · Around