2023

Style Over Substance: Evaluation Biases for Large Language Models

Wu, Minghao, Aji, Alham Fikri

Understand

As large language models (LLMs) continue to advance, accurately and comprehensively evaluating their performance becomes increasingly challenging.

  • Ranking the relative performance of LLMs based on Elo ratings, according to human judgment, is gaining more popularity.
  • However, the extent to which humans and LLMs are capable evaluators remains uncertain.
  • This study investigates the behavior of crowd-sourced and expert annotators, as well as LLMs, when comparing outputs from different models.

Reading the bibliography…