Fetching the paper…

ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation · Around