Fetching the paper…

Human Behavioral Benchmarking: Numeric Magnitude Comparison Effects in Large Language Models · Around