Fetching the paper…

On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification · Around