Fetching the paper…

Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning · Around