Fetching the paper…

MR-GSM8K: A Meta-Reasoning Benchmark for Large Language Model Evaluation · Around