Fetching the paper…

SPO: Multi-Dimensional Preference Sequential Alignment With Implicit Reward Modeling · Around