Fetching the paper…

Efficient Vision-Language Pretraining with Visual Concepts and Hierarchical Alignment · Around