Fetching the paper…

Scaling Pre-training to One Hundred Billion Data for Vision Language Models · Around