Fetching the paper…

Beyond Sparse Rewards: Enhancing Reinforcement Learning with Language Model Critique in Text Generation · Around