Fetching the paper…

Improving LLM General Preference Alignment via Optimistic Online Mirror Descent · Around