Fetching the paper…

Vision-Language Intelligence: Tasks, Representation Learning, and Large Models · Around