Fetching the paper…

DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning · Around