2024

The Fault in our Stars: Quality Assessment of Code Generation Benchmarks

Siddiq, Mohammed Latif, Dristi, Simantika, Saha, Joy et al.

Understand

Large Language Models (LLMs) are gaining popularity among software engineers.

  • A crucial aspect of developing effective code generation LLMs is to evaluate these models using a robust benchmark.
  • Evaluation benchmarks with quality issues can provide a false sense of performance.
  • In this work, we conduct the first-of-its-kind study of the quality of prompts within benchmarks used to compare the performance of different code generation models.

Reading the bibliography…