2023

UltraFeedback: Boosting Language Models with Scaled AI Feedback

Cui, Ganqu, Yuan, Lifan, Ding, Ning et al.

Understand

Learning from human feedback has become a pivot technique in aligning large language models (LLMs) with human preferences.

  • However, acquiring vast and premium human feedback is bottlenecked by time, labor, and human capability, resulting in small sizes or limited topics of current datasets.
  • This further hinders feedback learning as well as alignment research within the open-source community.
  • To address this issue, we explore how to go beyond human feedback and collect high-quality \textit{AI feedback} automatically for a scalable alternative.

Reading the bibliography…