Fetching the paper…

Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy Competition · Around