Fetching the paper…
Reading the bibliography…
In this paper, we present a challenging code reasoning task: vulnerability detection.
Measuring nominal scale agreement among many raters
Joseph L. Fleiss · 1971
Earlier work this paper cites.
The Balanced Accuracy and Its Posterior Distribution
Kay Henning Brodersen, Cheng Soon Ong, Klaas Enno Stephan, and Joachim M. Buhmann · 2010
Earlier work this paper cites.
Infer: An automatic program verifier for memory safety of c programs
Cristiano Calcagno and Dino Distefano · 2011
Earlier work this paper cites.
CVE-2017-9211, 2024
The MITRE Corporation · 2017
Earlier work this paper cites.
CVE-2018-16435, 2024
MITRE · 2018
Earlier work this paper cites.
A comprehensive study on deep learning bug characteristics
Md Johirul Islam, Giang Nguyen, Rangeet Pan, and Hridesh Rajan · 2019
Earlier work this paper cites.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2019
Earlier work this paper cites.
Devign: Effective vulnerability identification by learning comprehensive program semantics via graph neural networks
Yaqin Zhou, Shangqing Liu, Jingkai Siow, Xiaoning Du, and Yang Liu · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, et al · 2020
Earlier work this paper cites.
A C/C++ Code Vulnerability Dataset with Code Changes and CVE Summaries
Jiahao Fan, Yi Li, Shaohua Wang, and Tien N. Nguyen · 2020
Earlier work this paper cites.
CodeBERT: A pre-trained model for programming and natural languages
Zhangyin Feng, Daya Guo, Duyu Tang, et al · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, et al · 2020
Earlier work this paper cites.
Deep learning based vulnerability detection: Are we there yet?
S. Chakraborty, R. Krishna, Y. Ding, and B. Ray · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
The coding manual for qualitative researchers
Johnny Saldaña · 2021
Earlier work this paper cites.
D2A: A dataset built for ai-based vulnerability detection methods using differential analysis
Yunhui Zheng, Saurabh Pujar, Burn Lewis, et al · 2021
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe · 2022
Earlier work this paper cites.
How does in-context learning work? A framework for understanding the differences from traditional supervised learning
Sang Michael Xie and Sewon Min · 2022
Earlier work this paper cites.
Transformer-based vulnerability detection in code at edittime: Zero-shot, few-shot, or fine-tuning?
Aaron Chan, Anant Kharkar, and Roshanak Zilouchian others Moghaddam · 2023
Earlier work this paper cites.
Diversevul: A new vulnerable source code dataset for deep learning based vulnerability detection
Yizheng Chen, Zhoujie Ding, Lamya Alowain, Xinyun Chen, and David Wagner · 2023
Earlier work this paper cites.
Data quality for software vulnerability datasets
Roland Croft, M Ali Babar, and M Mehdi Kholoosi · 2023
Cited alongside, same era.
ChatGPT for vulnerability detection, classification, and repair: How far are we?
Michael Fu, Chakkrit Tantithamthavorn, Van Nguyen, and Trung Le · 2023
Cited alongside, same era.
How far have we gone in vulnerability detection using large language models, 2023
Zeyu Gao, Hao Wang, Yuchen Zhou, Wenyu Zhu, and Chao Zhang · 2023
Cited alongside, same era.
Gemini: A family of highly capable multimodal models
Gemini Team · 2023
Cited alongside, same era.
Large language models for code: Security hardening and adversarial testing
Jingxuan He and Martin Vechev · 2023
Cited alongside, same era.
DeepSeek-R1-Lite-Preview is now live: unleashing supercharged reasoning power!
DeepSeek · 2024
Closest in time.
Infer Static Analyzer, 2024
Facebook · 2024
Closest in time.
Prompting is all you need: Automated android bug replay with large language models
Sidong Feng and Chunyang Chen · 2024
Closest in time.
CRUXEval: A benchmark for code reasoning, understanding and execution, 2024
Alex Gu, Baptiste Rozière, Hugh Leather, Armando Solar-Lezama, Gabriel Synnaeve, and Sida I. Wang · 2024
Closest in time.
Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Y. Wu, Y. K. Li, Fuli Luo, Yingfei Xiong, and Wenfeng Liang · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, et al · 2023
Cited alongside, same era.
Understanding the effectiveness of large language models in detecting security vulnerabilities
Avishree Khare, Saikat Dutta, Ziyang Li, Alaia Solko-Breslin, Rajeev Alur, and Mayur Naik · 2023
Cited alongside, same era.
WizardCoder: Empowering code large language models with evol-instruct
Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang · 2023
Cited alongside, same era.
gpt-3.5-turbo-0613 announcement, June 2023
OpenAI · 2023
Cited alongside, same era.
HumanEval Benchmark, 2023
PapersWithCode · 2023
Cited alongside, same era.
Software vulnerability detection using large language models
M. Purba, A. Ghosh, B. J. Radford, and B. Chu · 2023
Cited alongside, same era.
Code Llama: Open foundation models for code
Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, et al · 2023
Cited alongside, same era.
HuggingFaceH4/starchat2-15b-v0.1
HuggingFaceH4 Team · 2024
Closest in time.
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, et al · 2024
Closest in time.
Starcoder 2 and the stack v2: The next generation, 2024
Anton Lozhkov, Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy-Poirier, Nouamane Tazi, Ao Tang, Dmytro Pykhtar, Jiawei Liu, Yuxiang Wei, Tianyang Liu, Max Tian, Denis Kocetkov, Arthur Zucker, Younes Belkada, Zijian Wang, Qian Liu, Dmitry Abulkhanov, Indraneil Paul, Zhuang Li, Wen-Ding Li, Megan Risdal, Jia Li, Jian Zhu, Terry Yue Zhuo, Evgenii Zheltonozhskii, Nii Osae Osae Dade, Wenhao Yu, Lucas Krauß, Naman Jain, Yixuan Su, Xuanli He, Manan Dey, Edoardo Abati, Yekun Chai, Niklas Muennighoff, Xiangru Tang, Muhtasham Oblokulov, Christopher Akiki, Marc Marone, Chenghao Mou, Mayank Mishra, Alex Gu, Binyuan Hui, Tri Dao, Armel Zebaze, Olivier Dehaene, Nicolas Patry, Canwen Xu, Julian McAuley, Han Hu, Torsten Scholak, Sebastien Paquet, Jennifer Robinson, Carolyn Jane Anderson, Nicolas Chapados, Mostofa Patwary, Nima Tajbakhsh, Yacine Jernite, Carlos Muñoz Ferrandis, Lingming Zhang, Sean Hughes, Thomas Wolf, Arjun Guha, Leandro von Werra, and Harm de Vries · 2024
Closest in time.
CWE - Common Weakness Enumeration
MITRE · 2024
Closest in time.
Yu Nong, Mohammed Aldeen, Long Cheng, Hongxin Hu, Feng Chen, and Haipeng Cai · 2024
Closest in time.
GPT-4 technical report, 2024
OpenAI · 2024
Closest in time.
Learning to reason with llms
OpenAI · 2024
Closest in time.
Finetuning large language models for vulnerability detection
Alexey Shestov, Anton Cheshkov, Rodion Levichev, et al · 2024
Closest in time.
GPTScan: Detecting logic vulnerabilities in smart contracts by combining gpt with program analysis
Yuqiang Sun, Daoyuan Wu, Yue Xue, et al · 2024
Closest in time.
Natural experiment
Wikipedia · 2024
Closest in time.
Large language models for test-free fault localization
Aidan ZH Yang, Claire Le Goues, Ruben Martins, and Vincent Hellendoorn · 2024
Closest in time.
Security Code Review by LLMs: A Deep Dive into Responses, 2024
Jiaxin Yu, Peng Liang, Yujia Fu, et al · 2024
Closest in time.
Imam Nur Bani Yusuf and Lingxiao Jiang · 2024
Closest in time.
Large language model for vulnerability detection: Emerging results and future directions
Xin Zhou, Ting Zhang, and David Lo · 2024
Closest in time.
Vulnerability detection with code language models: How far are we?
Yangruibo Ding, Yanjun Fu, Omniyyah Ibrahim, Chawin Sitawarin, Xinyun Chen, Basel Alomair, David Wagner, Baishakhi Ray, and Yizheng Chen · 2025
Closest in time.