2025

Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language Models

Tan, Huajie, Ji, Yuheng, Hao, Xiaoshuai et al.

Understand

Visual reasoning abilities play a crucial role in understanding complex multimodal data, advancing both domain-specific applications and artificial general intelligence (AGI).

  • Existing methods enhance Vision-Language Models (VLMs) through Chain-of-Thought (CoT) supervised fine-tuning using meticulously annotated data.
  • However, this approach may lead to overfitting and cognitive rigidity, limiting the model's generalization ability under domain shifts and reducing real-world applicability.
  • To overcome these limitations, we propose Reason-RFT, a two-stage reinforcement fine-tuning framework for visual reasoning.

Reading the bibliography…