Fetching the paper…

Angles Don't Lie: Unlocking Training-Efficient RL Through the Model's Own Signals · Around