Fetching the paper…
Reading the bibliography…
Existing benchmarks for evaluating the security risks and capabilities (e.g., vulnerability detection) of code-generating large language models (LLMs) face several key limitations: (1) limited coverage of risk and capabilities; (2) reliance on static evaluation metrics such as LLM judgments or rule-based detection, which lack the precision of dynamic analysis; and (3) a trade-off between data quality and benchmark scale.
Clangd, 2007
LLVM · 2007
Earlier work this paper cites.
libfuzzer, 2015
LLVM · 2015
Earlier work this paper cites.
Eed: Extended edit distance measure for machine translation
Peter Stanchev, Weiyue Wang, and Hermann Ney · 2019
Earlier work this paper cites.
Ac/c++ code vulnerability dataset with code changes and cve summaries
Jiahao Fan, Yi Li, Shaohua Wang, and Tien N Nguyen · 2020
Earlier work this paper cites.
Afl++: combining incremental steps of fuzzing research
Andrea Fioraldi, Dominik Maier, Heiko Eißfeldt, and Marc Heuse · 2020
Earlier work this paper cites.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Hackthebox
HackTheBox · 2021
Earlier work this paper cites.
Javaparser, 2021
JavaParser.org · 2021
Earlier work this paper cites.
Leetcode dataset
DeepSeek · 2022
Earlier work this paper cites.
Asleep at the keyboard? assessing the security of github copilot’s code contributions
Hammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt, and Ramesh Karri · 2022
Earlier work this paper cites.
Securityeval dataset: Mining vulnerability examples to evaluate machine learning-based code generation techniques
Mohammed Latif Siddiq and Joanna C. S. Santos · 2022
Earlier work this paper cites.
Securityeval dataset: mining vulnerability examples to evaluate machine learning-based code generation techniques
Mohammed Latif Siddiq and Joanna CS Santos · 2022
Earlier work this paper cites.
Purple llama cyberseceval: A secure coding benchmark for language models
Manish Bhatt, Sahana Chennabasappa, Cyrus Nikolaidis, Shengye Wan, Ivan Evtimov, Dominik Gabi, Daniel Song, Faizan Ahmad, Cornelius Aschermann, Lorenzo Fontana, et al · 2023
Earlier work this paper cites.
Codescore: Evaluating code generation by learning code execution
Yihong Dong, Jiazheng Ding, Xue Jiang, Ge Li, Zhuo Li, and Zhi Jin · 2023
Earlier work this paper cites.
Large language models for code: Security hardening and adversarial testing
Jingxuan He and Martin Vechev · 2023
Earlier work this paper cites.
Swe-bench: Can language models resolve real-world github issues?
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan · 2023
Earlier work this paper cites.
Ds-1000: A natural and reliable benchmark for data science code generation
Yuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang, Ruiqi Zhong, Luke Zettlemoyer, Wen-tau Yih, Daniel Fried, Sida Wang, and Tao Yu · 2023
Earlier work this paper cites.
Juliet java test suite, 2023
NIST · 2023
Earlier work this paper cites.
The formai dataset: Generative ai in software security through the lens of formal verification
Norbert Tihanyi, Tamas Bisztray, Ridhi Jain, Mohamed Amine Ferrag, Lucas C. Cordeiro, and Vasileios Mavroeidis · 2023
Earlier work this paper cites.
Llmseceval: A dataset of natural language prompts for security evaluations
Catherine Tony, Markus Mutas, Nicolás E Díaz Ferreyra, and Riccardo Scandariato · 2023
Earlier work this paper cites.
Deceptprompt: Exploiting llm-driven code generation via adversarial natural language instructions
Fangzhou Wu, Xiaogeng Liu, and Chaowei Xiao · 2023
Cited alongside, same era.
Cyberseceval 2: A wide-ranging cybersecurity evaluation suite for large language models
Manish Bhatt, Sahana Chennabasappa, Yue Li, Cyrus Nikolaidis, Daniel Song, Shengye Wan, Faizan Ahmad, Cornelius Aschermann, Yaohui Chen, Dhaval Kapil, et al · 2024
Cited alongside, same era.
An empirical study of static analysis tools for secure code review
Wachiraphan Charoenwet, Patanamon Thongtanunam, Van-Thuan Pham, and Christoph Treude · 2024
Cited alongside, same era.
Rmcbench: Benchmarking large language models’ resistance to malicious code
Jiachi Chen, Qingyuan Zhong, Yanlin Wang, Kaiwen Ning, Yongkun Liu, Zenan Xu, Zhe Zhao, Ting Chen, and Zibin Zheng · 2024
Cited alongside, same era.
Mitre caldera: A scalable, automated adversary emulation platform
Trustllm: Trustworthiness in large language models
Lichao Sun, Yue Huang, Haoran Wang, Siyuan Wu, Qihui Zhang, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, et al · 2024
Closest in time.
Judging the judges: Evaluating alignment and vulnerabilities in llms-as-judges
Aman Singh Thakur, Kartik Choudhary, Venkat Srinik Ramayapally, Sankaran Vaidyanathan, and Dieuwke Hupkes · 2024
Closest in time.
Debugbench: Evaluating debugging capability of large language models
Runchu Tian, Yining Ye, Yujia Qin, Xin Cong, Yankai Lin, Zhiyuan Liu, and Maosong Sun · 2024
Closest in time.
Cybermetric: A benchmark dataset for evaluating large language models knowledge in cybersecurity
Norbert Tihanyi, Mohamed Amine Ferrag, Ridhi Jain, and Merouane Debbah · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
MITRE Corporation · 2024
Cited alongside, same era.
Vulnerability detection with code language models: How far are we?
Yangruibo Ding, Yanjun Fu, Omniyyah Ibrahim, Chawin Sitawarin, Xinyun Chen, Basel Alomair, David Wagner, Baishakhi Ray, and Yizheng Chen · 2024
Cited alongside, same era.
Cycle: Learning to self-refine the code generation
Yangruibo Ding, Marcus J Min, Gail Kaiser, and Baishakhi Ray · 2024
Cited alongside, same era.
Constrained decoding for secure code generation
Yanjun Fu, Ethan Baker, Yu Ding, and Yizheng Chen · 2024
Cited alongside, same era.
Autopenbench: Benchmarking generative agents for penetration testing
Luca Gioacchini, Marco Mellia, Idilio Drago, Alexander Delsanto, Giuseppe Siracusano, and Roberto Bifulco · 2024
Cited alongside, same era.
Cruxeval: A benchmark for code reasoning, understanding and execution
Alex Gu, Baptiste Rozière, Hugh Leather, Armando Solar-Lezama, Gabriel Synnaeve, and Sida I Wang · 2024
Cited alongside, same era.
Redcode: Risky code execution and generation benchmark for code agents
Chengquan Guo, Xun Liu, Chulin Xie, Andy Zhou, Yi Zeng, Zinan Lin, Dawn Song, and Bo Li · 2024
Cited alongside, same era.
Codelmsec benchmark: Systematically evaluating and finding security vulnerabilities in black-box code language models
Hossein Hajipour, Keno Hassler, Thorsten Holz, Lea Schönherr, and Mario Fritz · 2024
Cited alongside, same era.
Together AI · 2024
Closest in time.
Llms cannot reliably identify and reason about security vulnerabilities (yet?): A comprehensive evaluation, framework, and benchmarks
Saad Ullah, Mingji Han, Saurabh Pujar, Hammond Pearce, Ayse Coskun, and Gianluca Stringhini · 2024
Closest in time.
Shengye Wan, Cyrus Nikolaidis, Daniel Song, David Molnar, James Crnkovich, Jayson Grace, Manish Bhatt, Sahana Chennabasappa, Spencer Whitman, Stephanie Ding, et al · 2024
Closest in time.
Intercode: Standardizing and benchmarking interactive coding with execution feedback
John Yang, Akshara Prabhakar, Karthik Narasimhan, and Shunyu Yao · 2024
Closest in time.
R-judge: Benchmarking safety risk awareness for llm agents
Tongxin Yuan, Zhiwei He, Lingzhong Dong, Yiming Wang, Ruijie Zhao, Tian Xia, Lizhen Xu, Binglin Zhou, Fangqi Li, Zhuosheng Zhang, Rui Wang, and Gongshen Liu · 2024
Closest in time.
Cybench: A framework for evaluating cybersecurity capabilities and risk of language models
Andy K Zhang, Neil Perry, Riya Dulepet, Eliot Jones, Justin W Lin, Joey Ji, Celeste Menders, Gashon Hussein, Samantha Liu, Donovan Jasper, et al · 2024
Closest in time.
Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions
Terry Yue Zhuo, Minh Chien Vu, Jenny Chim, Han Hu, Wenhao Yu, Ratnadira Widyasari, Imam Nur Bani Yusuf, Haolan Zhan, Junda He, Indraneil Paul, et al · 2024
Closest in time.
Claude 3.7 sonnet and claude code, February 2025
Anthropic · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
DeepSeek-AI · 2025
Closest in time.
Cyberseceval 4: Advancing the evaluation of cybersecurity risks and capabilities in large language models, 2025
Meta · 2025
Closest in time.
Hello gpt-4o, May 2024
OpenAI · 2025
Closest in time.
Introducing openai o3 and o4-mini
OpenAI · 2025
Closest in time.
Cweval: Outcome-driven evaluation on functionality and security of llm code generation
Jinjun Peng, Leyi Cui, Kele Huang, Junfeng Yang, and Baishakhi Ray · 2025
Closest in time.
Qwen · 2025
Closest in time.
Qwq-32b: Embracing the power of reinforcement learning
Qwen · 2025
Closest in time.
Baxbench: Can llms generate correct and secure backends?
Mark Vero, Niels Mündler, Victor Chibotaru, Veselin Raychev, Maximilian Baader, Nikola Jovanović, Jingxuan He, and Martin Vechev · 2025
Closest in time.