Fetching the paper…

Adversarial Preference Optimization: Enhancing Your Alignment via RM-LLM Game · Around