2024

LLM Comparator: Visual Analytics for Side-by-Side Evaluation of Large Language Models

Kahng, Minsuk, Tenney, Ian, Pushkarna, Mahima et al.

Understand

Automatic side-by-side evaluation has emerged as a promising approach to evaluating the quality of responses from large language models (LLMs).

  • However, analyzing the results from this evaluation approach raises scalability and interpretability challenges.
  • In this paper, we present LLM Comparator, a novel visual analytics tool for interactively analyzing results from automatic side-by-side evaluation.
  • The tool supports interactive workflows for users to understand when and why a model performs better or worse than a baseline model, and how the responses from two models are qualitatively different.

Reading the bibliography…