2024

TurkishMMLU: Measuring Massive Multitask Language Understanding in Turkish

Yüksel, Arda, Köksal, Abdullatif, Şenel, Lütfi Kerem et al.

Understand

Multiple choice question answering tasks evaluate the reasoning, comprehension, and mathematical abilities of Large Language Models (LLMs).

  • While existing benchmarks employ automatic translation for multilingual evaluation, this approach is error-prone and potentially introduces culturally biased questions, especially in social sciences.
  • We introduce the first multitask, multiple-choice Turkish QA benchmark, TurkishMMLU, to evaluate LLMs' understanding of the Turkish language.
  • TurkishMMLU includes over 10,000 questions, covering 9 different subjects from Turkish high-school education curricula.

Reading the bibliography…