Fetching the paper…
Reading the bibliography…
Long contexts of recent LLMs have enabled a new use case: asking models to find security vulnerabilities in entire codebases.
“OSS-Fuzz - Google’s continuous fuzzing service for open source software”
Kostya Serebryany · 2017
Earlier work this paper cites.
Yizheng Chen et al · 2023
Earlier work this paper cites.
“A python library for confidence intervals”
Jacob Gildenblat · 2023
Earlier work this paper cites.
“CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models”, 2024
Manish Bhatt et al · 2024
Earlier work this paper cites.
“Project Naptime: Evaluating Offensive Security Capabilities of Large Language Models”, 2024
Sergei Glazunov and Mark Brand · 2024
Cited alongside, same era.
“OSS-Fuzz-Gen: Automated Fuzz Target Generation”, 2024
Dongge Liu et al · 2024
Cited alongside, same era.
“Evaluating Frontier Models for Dangerous Capabilities”, 2024
Mary Phuong et al · 2024
Cited alongside, same era.
“An Empirical Evaluation of LLMs for Solving Offensive Security Challenges”, 2024
Minghao Shao et al · 2024
Closest in time.
Minghao Shao et al · 2024
Closest in time.
“VulEval: Towards Repository-Level Evaluation of Software Vulnerability Detection”, 2024
Xin-Cheng Wen et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…