2024

Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Chiang, Wei-Lin, Zheng, Lianmin, Sheng, Ying et al.

Understand

Large Language Models (LLMs) have unlocked new capabilities and applications; however, evaluating the alignment with human preferences still poses significant challenges.

  • To address this issue, we introduce Chatbot Arena, an open platform for evaluating LLMs based on human preferences.
  • Our methodology employs a pairwise comparison approach and leverages input from a diverse user base through crowdsourcing.
  • The platform has been operational for several months, amassing over 240K votes.

Reading the bibliography…