2025

Exploring Data Scaling Trends and Effects in Reinforcement Learning from Human Feedback

Shen, Wei, Liu, Guanlin, Wu, Zheng et al.

Understand

Reinforcement Learning from Human Feedback (RLHF) is crucial for aligning large language models with human preferences.

  • While recent research has focused on algorithmic improvements, the importance of prompt-data construction has been overlooked.
  • This paper addresses this gap by exploring data-driven bottlenecks in RLHF performance scaling, particularly reward hacking and decreasing response diversity.
  • We introduce a hybrid reward system combining reasoning task verifiers (RTV) and a generative reward model (GenRM) to mitigate reward hacking.

Reading the bibliography…