Fetching the paper…

Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment · Around