Fetching the paper…

InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning · Around