2024

Khayyam Challenge (PersianMMLU): Is Your LLM Truly Wise to The Persian Language?

Ghahroodi, Omid, Nouri, Marzia, Sanian, Mohammad Vali et al.

Understand

Evaluating Large Language Models (LLMs) is challenging due to their generative nature, necessitating precise evaluation methodologies.

  • Additionally, non-English LLM evaluation lags behind English, resulting in the absence or weakness of LLMs for many languages.
  • In response to this necessity, we introduce Khayyam Challenge (also known as PersianMMLU), a meticulously curated collection comprising 20,192 four-choice questions sourced from 38 diverse tasks extracted from Persian examinations, spanning a wide spectrum of subjects, complexities, and ages.
  • The primary objective of the Khayyam Challenge is to facilitate the rigorous evaluation of LLMs that support the Persian language.

Reading the bibliography…