2023

Improving Generalization of Alignment with Human Preferences through Group Invariant Learning

Zheng, Rui, Shen, Wei, Hua, Yuan et al.

Understand

The success of AI assistants based on language models (LLMs) hinges crucially on Reinforcement Learning from Human Feedback (RLHF), which enables the generation of responses more aligned with human preferences.

  • As universal AI assistants, there's a growing expectation for them to perform consistently across various domains.
  • However, previous work shows that Reinforcement Learning (RL) often exploits shortcuts to attain high rewards and overlooks challenging samples.
  • This focus on quick reward gains undermines both the stability in training and the model's ability to generalize to new, unseen data.

Reading the bibliography…