Fetching the paper…

PipelineRL: Faster On-policy Reinforcement Learning for Long Sequence Generation · Around