Fetching the paper…

Forward KL Regularized Preference Optimization for Aligning Diffusion Policies · Around