Fetching the paper…

CLR-Bench: Evaluating Large Language Models in College-level Reasoning · Around