Fetching the paper…

ZeroSumEval: An Extensible Framework For Scaling LLM Evaluation with Inter-Model Competition · Around