2025

J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning

Whitehouse, Chenxi, Wang, Tianlu, Yu, Ping et al.

Understand

The progress of AI is bottlenecked by the quality of evaluation, making powerful LLM-as-a-Judge models a core solution.

  • The efficacy of these judges depends on their chain-of-thought reasoning, creating a critical need for methods that can effectively optimize this reasoning process.
  • In this work, we introduce J1, a reinforcement learning framework for teaching LLM judges to think before making decisions.
  • Our core contribution lies in converting all judgment tasks for non-verifiable and verifiable prompts into a unified format with verifiable rewards, enabling direct optimization of evaluation quality while mitigating positional bias.

Reading the bibliography…