Fetching the paper…

ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models · Around