Fetching the paper…

CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification · Around