Fetching the paper…

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design · Around