2024

Benchmark Data Contamination of Large Language Models: A Survey

Xu, Cheng, Guan, Shuhao, Greene, Derek et al.

Understand

The rapid development of Large Language Models (LLMs) like GPT-4, Claude-3, and Gemini has transformed the field of natural language processing.

  • However, it has also resulted in a significant issue known as Benchmark Data Contamination (BDC).
  • This occurs when language models inadvertently incorporate evaluation benchmark information from their training data, leading to inaccurate or unreliable performance during the evaluation phase of the process.
  • This paper reviews the complex challenge of BDC in LLM evaluation and explores alternative assessment methods to mitigate the risks associated with traditional benchmarks.

Reading the bibliography…