2024

Multitask-based Evaluation of Open-Source LLM on Software Vulnerability

Yin, Xin, Ni, Chao, Wang, Shaohua

Understand

This paper proposes a pipeline for quantitatively evaluating interactive Large Language Models (LLMs) using publicly available datasets.

  • We carry out an extensive technical evaluation of LLMs using Big-Vul covering four different common software vulnerability tasks.
  • This evaluation assesses the multi-tasking capabilities of LLMs based on this dataset.
  • We find that the existing state-of-the-art approaches and pre-trained Language Models (LMs) are generally superior to LLMs in software vulnerability detection.

Reading the bibliography…