Fetching the paper…

Improving Reinforcement Learning from Human Feedback Using Contrastive Rewards · Around