2024

Does RLHF Scale? Exploring the Impacts From Data, Model, and Method

Hou, Zhenyu, Du, Pengfan, Niu, Yilin et al.

Understand

This study explores the scaling properties of Reinforcement Learning from Human Feedback (RLHF) in Large Language Models (LLMs).

  • Although RLHF is considered an important step in post-training of LLMs, its scaling potential is still largely unknown.
  • We systematically analyze key components in the RLHF framework--model size, data composition, and inference budget--and their impacts on performance.
  • Our findings show that increasing data diversity and volume improves reward model performance, helping process-supervision models scale better.

Reading the bibliography…