Fetching the paper…

InstructCV: Instruction-Tuned Text-to-Image Diffusion Models as Vision Generalists · Around