2024

CodeMind: Evaluating Large Language Models for Code Reasoning

Liu, Changshu, Chen, Yang, Jabbarvand, Reyhaneh

Understand

Large Language Models (LLMs) have been widely used to automate programming tasks.

  • Their capabilities have been evaluated by assessing the quality of generated code through tests or proofs.
  • The extent to which they can reason about code is a critical question revealing important insights about their true capabilities.
  • This paper introduces CodeMind, a framework designed to gauge the code reasoning abilities of LLMs through the following explicit and implicit code reasoning tasks: Independent Execution Reasoning (IER), Specification Reasoning (SR) and Dynamic Semantics Reasoning (DSR).

Built on

Nothing clear enough to list yet.

Similar

Nothing clear enough to list yet.

Then

Nothing clear enough to list yet.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…