Fetching the paper…

Unified Language-Vision Pretraining in LLM with Dynamic Discrete Visual Tokenization · Around