Fetching the paper…
Reading the bibliography…
We propose and release a new vulnerable source code dataset.
A few billion lines of code later: using static analysis to find bugs in the real world
Al Bessey, Ken Block, Ben Chelf, Andy Chou, Seth Hallem Bryan Fulton, Charles Henri-Gros, Asya Kamsky, Scott McPeak, and Dawson Engler. 2010 · 2010
Earlier work this paper cites.
Report on the static analysis tool exposition (sate) iv
Vadim Okun, Aurelien Delaitre, Paul E Black, et al · 2013
Earlier work this paper cites.
Modeling and discovering vulnerabilities with code property graphs. In 2014 IEEE Symposium on Security and Privacy . IEEE, 590–604
Fabian Yamaguchi, Nico Golde, Daniel Arp, and Konrad Rieck. 2014 · 2014
Earlier work this paper cites.
Gated graph sequence neural networks
Yujia Li, Daniel Tarlow, Marc Brockschmidt, and Richard Zemel. 2015 · 2015
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Vuldeepecker: A deep learning-based system for vulnerability detection
Zhen Li, Deqing Zou, Shouhuai Xu, Xinyu Ou, Hai Jin, Sujuan Wang, Zhijun Deng, and Yuyi Zhong. 2018 · 2018
Earlier work this paper cites.
Automated vulnerability detection in source code using deep representation learning. In 2018 17th IEEE international conference on machine learning and applications (ICMLA) . IEEE, 757–762
Rebecca Russell, Louis Kim, Lei Hamilton, Tomo Lazovich, Jacob Harer, Onur Ozdemir, Paul Ellingwood, and Marc McConley. 2018 · 2018
Earlier work this paper cites.
Codesearchnet challenge: Evaluating the state of semantic code search
Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt. 2019 · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Yaqin Zhou, Shangqing Liu, Jingkai Siow, Xiaoning Du, and Yang Liu. 2019 · 2019
Earlier work this paper cites.
AC/C++ code vulnerability dataset with code changes and CVE summaries. In Proceedings of the 17th International Conference on Mining Software Repositories . 508–512
Jiahao Fan, Yi Li, Shaohua Wang, and Tien N Nguyen. 2020 · 2020
Cited alongside, same era.
Codebert: A pre-trained model for programming and natural languages
Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, et al · 2020
Cited alongside, same era.
Graphcodebert: Pre-training code representations with data flow
Daya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng, Duyu Tang, Shujie Liu, Long Zhou, Nan Duan, Alexey Svyatkovskiy, Shengyu Fu, et al · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Cited alongside, same era.
Yue Wang, Weishi Wang, Shafiq Joty, and Steven CH Hoi. 2021b · 2021
Later among the works it cites.
D2A: a dataset built for AI-based vulnerability detection methods using differential analysis. In 2021 IEEE/ACM 43rd International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP) . IEEE, 111–120
Yunhui Zheng, Saurabh Pujar, Burn Lewis, Luca Buratti, Edward Epstein, Bo Yang, Jim Laredo, Alessandro Morari, and Zhong Su. 2021 · 2021
Later among the works it cites.
NatGen: generative pre-training by “naturalizing” source code. In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering . 18–30
Saikat Chakraborty, Toufique Ahmed, Yangruibo Ding, Premkumar T Devanbu, and Baishakhi Ray. 2022 · 2022
Later among the works it cites.
Building a Commit-level Dataset of Real-world Vulnerabilities. In Proceedings of the Twelveth ACM Conference on Data and Application Security and Privacy . 101–106
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
CVEfixes: automated collection of vulnerabilities and their fixes from open-source software. In Proceedings of the 17th International Conference on Predictive Models and Data Analytics in Software Engineering . 30–39
Guru Bhandari, Amara Naseer, and Leon Moonen. 2021 · 2021
Cited alongside, same era.
Deep learning based vulnerability detection: Are we there yet
Saikat Chakraborty, Rahul Krishna, Yangruibo Ding, and Baishakhi Ray. 2021 · 2021
Cited alongside, same era.
Sysevr: A framework for using deep learning to detect software vulnerabilities
Zhen Li, Deqing Zou, Shouhuai Xu, Hai Jin, Yawei Zhu, and Zhaoxuan Chen. 2021 · 2021
Cited alongside, same era.
Codexglue: A machine learning benchmark dataset for code understanding and generation
Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin Clement, Dawn Drain, Daxin Jiang, Duyu Tang, et al · 2021
Cited alongside, same era.
CrossVul: a cross-language vulnerability dataset with commit data. In Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering . 1565–1569
Georgios Nikitopoulos, Konstantina Dritsa, Panos Louridas, and Dimitris Mitropoulos. 2021 · 2021
Cited alongside, same era.
Patchdb: A large-scale security patch dataset. In 2021 51st Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN) . IEEE, 149–160
Xinda Wang, Shu Wang, Pengbin Feng, Kun Sun, and Sushil Jajodia. 2021a · 2021
Cited alongside, same era.
National Vulnerability Database
National Institute of Standards and Technology. Last accessed on March 19, 2023a
Cited in the paper.
NIST Software Assurance Reference Dataset
National Institute of Standards and Technology. Last accessed on March 19, 2023b
Cited in the paper.
Alexis Challande, Robin David, and Guénaël Renault. 2022 · 2022
Later among the works it cites.
2022 CWE Top 25 Most Dangerous Software Weaknesses
The MITRE Corporation. Last accessed on March 28, 2023 · 2022
Later among the works it cites.
Transformer-Based Language Models for Software Vulnerability Detection. In Proceedings of the 38th Annual Computer Security Applications Conference . 481–496
Chandra Thapa, Seung Ick Jang, Muhammad Ejaz Ahmed, Seyit Camtepe, Josef Pieprzyk, and Surya Nepal. 2022 · 2022
Later among the works it cites.
A Systematic Evaluation of Large Language Models of Code
Frank F Xu, Uri Alon, Graham Neubig, and Vincent J Hellendoorn. 2022 · 2022
Later among the works it cites.
Data quality for software vulnerability datasets. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE)
Roland Croft, M Ali Babar, and Mehdi Kholoosi. 2023 · 2023
Closest in time.
VulChecker: Graph-based Vulnerability Localization in Source Code. In USENIX Security 2023
Yisroel Mirsky, George Macon, Michael Brown, Carter Yagemann, Matthew Pruett, Evan Downing, Sukarno Mertoguno, and Wenke Lee. 2023 · 2023
Closest in time.
An empirical study of deep learning models for vulnerability detection. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE)
Benjamin Steenhoek, Md Mahbubur Rahman, Richard Jiles, and Wei Le. 2023 · 2023
Closest in time.