Fetching the paper…

Selective Preference Optimization via Token-Level Reward Function Estimation · Around