Fetching the paper…
Reading the bibliography…
Recent results of machine learning for automatic vulnerability detection (ML4VD) have been very promising.
Overfitting and undercomputing in machine learning
Tom Dietterich · 1995
Earlier work this paper cites.
On the naturalness of software
Abram Hindle, Earl T. Barr, Zhendong Su, Mark Gabel, and Premkumar Devanbu · 2012
Earlier work this paper cites.
The national vulnerability database (nvd): Overview, 2013-12-18 2013
Harold Booth, Doug Rike, and Gregory Witte · 2013
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
A software assurance reference dataset: Thousands of programs with known bugs, 2018-04-16 2018
Paul Black · 2018
Earlier work this paper cites.
VulDeePecker: A deep learning-based system for vulnerability detection
Zhen Li, Deqing Zou, Shouhuai Xu, Xinyu Ou, Hai Jin, Sujuan Wang, Zhijun Deng, and Yuyi Zhong · 2018
Earlier work this paper cites.
A manually-curated dataset of fixes to vulnerabilities of open-source software
Serena E. Ponta, Henrik Plate, Antonino Sabetta, Michele Bezzi, and Cédric Dangremont · 2019
Earlier work this paper cites.
An overview of overfitting and its solutions
Xue Ying · 2019
Earlier work this paper cites.
Devign: Effective Vulnerability Identification by Learning Comprehensive Program Semantics via Graph Neural Networks
Yaqin Zhou, Shangqing Liu, Jingkai Siow, Xiaoning Du, and Yang Liu · 2019
Earlier work this paper cites.
Adversarial robustness for code
Pavol Bielik and Martin Vechev · 2020
Earlier work this paper cites.
Codebert: A pre-trained model for programming and natural languages, 2020
Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, and Ming Zhou · 2020
Earlier work this paper cites.
Graphcodebert: Pre-training code representations with data flow
Daya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng, Duyu Tang, Shujie Liu, Long Zhou, Nan Duan, Alexey Svyatkovskiy, Shengyu Fu, et al · 2020
Earlier work this paper cites.
How often do single-statement bugs occur? the manysstubs4j dataset
Rafael-Michael Karampatsis and Charles Sutton · 2020
Earlier work this paper cites.
Adversarial examples for models of code
Noam Yefet, Uri Alon, and Eran Yahav · 2020
Earlier work this paper cites.
Generating adversarial examples for holding robustness of source code processing models
Huangzhao Zhang, Zhuo Li, Ge Li, L. Ma, Yang Liu, and Zhi Jin · 2020
Cited alongside, same era.
Unified pre-training for program understanding and generation
Wasi Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang · 2021
Cited alongside, same era.
Assessing robustness of ml-based program analysis tools using metamorphic program transformations
Leonhard Applis, Annibale Panichella, and Arie van Deursen · 2021
Cited alongside, same era.
Cvefixes: Automated collection of vulnerabilities and their fixes from open-source software
Guru Bhandari, Amara Naseer, and Leon Moonen · 2021
Cited alongside, same era.
Deep learning based vulnerability detection: Are we there yet?
Saikat Chakraborty, Rahul Krishna, Yangruibo Ding, and Baishakhi Ray · 2021
Cited alongside, same era.
Dos and don’ts of machine learning in computer security
Daniel Arp, Erwin Quiring, Feargus Pendlebury, Alexander Warnecke, Fabio Pierazzi, Christian Wressnegger, Lorenzo Cavallaro, and Konrad Rieck · 2022
Later among the works it cites.
Vul4j: A dataset of reproducible java vulnerabilities geared towards the study of program repair techniques
Quang-Cuong Bui, Riccardo Scandariato, and Nicolás E. Díaz Ferreyra · 2022
Later among the works it cites.
Unixcoder: Unified cross-modal pre-training for code representation
Daya Guo, Shuai Lu, Nan Duan, Yanlin Wang, Ming Zhou, and Jian Yin · 2022
Later among the works it cites.
Vulberta: Simplified source code pre-training for vulnerability detection
Hazim Hanif and Sergio Maffeis · 2022
Later among the works it cites.
On distribution shift in learning-based bug detectors
Jingxuan He, Luca Beurer-Kellner, and Martin Vechev · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, Dawn Song, Jacob Steinhardt, and Justin Gilmer · 2021
Cited alongside, same era.
Towards making deep learning-based vulnerability detectors robust, 2021
Zhen Li, Jing Tang, Deqing Zou, Qian Chen, Shouhuai Xu, Chao Zhang, Yichen Li, and Hai Jin · 2021
Cited alongside, same era.
Codexglue: A machine learning benchmark dataset for code understanding and generation
Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin Clement, Dawn Drain, Daxin Jiang, Duyu Tang, Ge Li, Lidong Zhou, Linjun Shou, Long Zhou, Michele Tufano, MING GONG, Ming Zhou, Nan Duan, Neel Sundaresan, Shao Kun Deng, Shengyu Fu, and Shujie LIU · 2021
Cited alongside, same era.
https://microsoft.github.io/CodeXGLUE/#LB-DefectDetection
Codexglue leaderboards, 2021 · 2021
Cited alongside, same era.
CoTexT: Multi-task learning with code-text transformer
Long Phan, Hieu Tran, Daniel Le, Hieu Nguyen, James Annibal, Alec Peltekian, and Yanfang Ye · 2021
Cited alongside, same era.
On the generalizability of neural program models with respect to semantic-preserving program transformations
Md Rafiqul Islam Rabin, Nghi D.Q. Bui, Ke Wang, Yijun Yu, Lingxiao Jiang, and Mohammad Amin Alipour · 2021
Cited alongside, same era.
Toward causal representation learning
B. Schölkopf, F. Locatello, S. Bauer, N. R. Ke, N. Kalchbrenner, A. Goyal, and Y. Bengio · 2021
Cited alongside, same era.
Semantic robustness of models of source code
Jordan Henkel, Goutham Ramakrishnan, Zi Wang, Aws Albarghouthi, Somesh Jha, and Thomas Reps · 2022
Later among the works it cites.
A closer look into transformer-based code intelligence through code transformation: Challenges and opportunities, 2022
Yaoxian Li, Shiyi Qi, Cuiyun Gao, Yun Peng, David Lo, Zenglin Xu, and Michael R. Lyu · 2022
Later among the works it cites.
https://github.com/ICL-ml4csec/VulBERTa/tree/main/data
Vuldeepecker function-level dataset, 2022 · 2022
Later among the works it cites.
Multipas: applying program transformations to introductory programming assignments for data augmentation
Pedro Orvalho, Mikoláš Janota, and Vasco Manquinho · 2022
Later among the works it cites.
Natural attack for pre-trained models of code
Zhou Yang, Jieke Shi, Junda He, and David Lo · 2022
Later among the works it cites.
Towards robustness of deep program processing models—detection, estimation, and enhancement
Huangzhao Zhang, Zhiyi Fu, Ge Li, Lei Ma, Zhehao Zhao, Hua’an Yang, Yizhe Sun, Yang Liu, and Zhi Jin · 2022
Later among the works it cites.
Diversevul: A new vulnerable source code dataset for deep learning based vulnerability detection
Yizheng Chen, Zhoujie Ding, Lamya Alowain, Xinyun Chen, and David Wagner · 2023
Closest in time.
An empirical study of deep learning models for vulnerability detection
Benjamin Steenhoek, Md Mahbubur Rahman, Richard Jiles, and Wei Le · 2023
Closest in time.
Towards causal deep learning for vulnerability detection
Md Mahbubur Rahman, Ira Ceka, Chengzhi Mao, Saikat Chakraborty, Baishakhi Ray, and Wei Le · 2024
Closest in time.