Fetching the paper…

Preference as Reward, Maximum Preference Optimization with Importance Sampling · Around