Fetching the paper…

ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning · Around