Fetching the paper…

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment · Around