Fetching the paper…
Reading the bibliography…
In the context of the rising interest in code language models (code LMs) and vulnerability detection, we study the effectiveness of code LMs for detecting vulnerabilities.
Report on the static analysis tool exposition (sate) iv
Vadim Okun, Aurelien Delaitre, Paul E Black, et al · 2013
Earlier work this paper cites.
Learning natural coding conventions
Miltiadis Allamanis, Earl T. Barr, Christian Bird, and Charles Sutton · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Vuldeepecker: A deep learning-based system for vulnerability detection
Zhen Li, Deqing Zou, Shouhuai Xu, Xinyu Ou, Hai Jin, Sujuan Wang, Zhijun Deng, and Yuyi Zhong · 2018
Earlier work this paper cites.
A survey of machine learning for big code and naturalness
Miltiadis Allamanis, Earl T. Barr, Premkumar Devanbu, and Charles Sutton · 2018
Earlier work this paper cites.
Devign: Effective vulnerability identification by learning comprehensive program semantics via graph neural networks
Yaqin Zhou, Shangqing Liu, Jingkai Siow, Xiaoning Du, and Yang Liu · 2019
Earlier work this paper cites.
Codexglue – defect detection, 2019
Microsoft · 2019
Earlier work this paper cites.
The adverse effects of code duplication in machine learning models of code
Miltiadis Allamanis · 2019
Earlier work this paper cites.
Survey on deep learning with class imbalance
Justin M. Johnson and Taghi M. Khoshgoftaar · 2019
Earlier work this paper cites.
Ac/c++ code vulnerability dataset with code changes and cve summaries
Jiahao Fan, Yi Li, Shaohua Wang, and Tien N Nguyen · 2020
Earlier work this paper cites.
Codebert: A pre-trained model for programming and natural languages
Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, et al · 2020
Earlier work this paper cites.
Graphcodebert: Pre-training code representations with data flow
Daya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng, Duyu Tang, Shujie Liu, Long Zhou, Nan Duan, Alexey Svyatkovskiy, Shengyu Fu, et al · 2020
Earlier work this paper cites.
Github copilot: Your ai pair programmer
GitHub · 2021
Earlier work this paper cites.
Deep learning based vulnerability detection: Are we there yet
Saikat Chakraborty, Rahul Krishna, Yangruibo Ding, and Baishakhi Ray · 2021
Earlier work this paper cites.
Codexglue: A machine learning benchmark dataset for code understanding and generation
Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin Clement, Dawn Drain, Daxin Jiang, Duyu Tang, et al · 2021
Earlier work this paper cites.
Cvefixes: automated collection of vulnerabilities and their fixes from open-source software
Guru Bhandari, Amara Naseer, and Leon Moonen · 2021
Earlier work this paper cites.
Crossvul: a cross-language vulnerability dataset with commit data
Georgios Nikitopoulos, Konstantina Dritsa, Panos Louridas, and Dimitris Mitropoulos · 2021
Earlier work this paper cites.
Sysevr: A framework for using deep learning to detect software vulnerabilities
Zhen Li, Deqing Zou, Shouhuai Xu, Hai Jin, Yawei Zhu, and Zhaoxuan Chen · 2021
Cited alongside, same era.
Confident learning: Estimating uncertainty in dataset labels
Curtis Northcutt, Lu Jiang, and Isaac Chuang · 2021
Cited alongside, same era.
Yue Wang, Weishi Wang, Shafiq Joty, and Steven CH Hoi · 2021
Cited alongside, same era.
SimCSE: Simple contrastive learning of sentence embeddings
Tianyu Gao, Xingcheng Yao, and Danqi Chen · 2021
Cited alongside, same era.
Unified pre-training for program understanding and generation
Wasi Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang · 2021
Cited alongside, same era.
Do language models learn semantics of code? a case study in vulnerability detection, 2023
Benjamin Steenhoek, Md Mahbubur Rahman, Shaila Sharmin, and Wei Le · 2023
Later among the works it cites.
Memorization and generalization in neural code intelligence models
Md Rafiqul Islam Rabin, Aftab Hussain, Mohammad Amin Alipour, and Vincent J. Hellendoorn · 2023
Later among the works it cites.
Concord: Clone-aware contrastive learning for source code
Yangruibo Ding, Saikat Chakraborty, Luca Buratti, Saurabh Pujar, Alessandro Morari, Gail Kaiser, and Baishakhi Ray · 2023
Later among the works it cites.
Codegen2: Lessons for training llms on programming and natural languages, 2023
Erik Nijkamp, Hiroaki Hayashi, Caiming Xiong, Silvio Savarese, and Yingbo Zhou · 2023
Later among the works it cites.
Understanding the effectiveness of large language models in detecting security vulnerabilities
Avishree Khare, Saikat Dutta, Ziyang Li, Alaia Solko-Breslin, Rajeev Alur, and Mayur Naik · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Michael Fu and Chakkrit Tantithamthavorn · 2022
Cited alongside, same era.
UniXcoder: Unified cross-modal pre-training for code representation
Daya Guo, Shuai Lu, Nan Duan, Yanlin Wang, Ming Zhou, and Jian Yin · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback, 2022
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe · 2022
Cited alongside, same era.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou · 2022
Cited alongside, same era.
Amazon codewhisperer: Build applications faster and more securely with your ai coding companion
Amazon · 2023
Cited alongside, same era.
Diversevul: A new vulnerable source code dataset for deep learning based vulnerability detection
Yizheng Chen, Zhoujie Ding, Lamya Alowain, Xinyun Chen, and David Wagner · 2023
Cited alongside, same era.
An empirical study of deep learning models for vulnerability detection
Benjamin Steenhoek, Md Mahbubur Rahman, Richard Jiles, and Wei Le · 2023
Cited alongside, same era.
Transactions on Machine Learning Research
Starcoder: may the source be with you! · 2023
Later among the works it cites.
Can large language models identify and reason about security vulnerabilities? not yet
Saad Ullah, Mingji Han, Saurabh Pujar, Hammond Pearce, Ayse Coskun, and Gianluca Stringhini · 2023
Later among the works it cites.
The larger they are, the harder they fail: Language models do not recognize identifier swaps in python
Antonio Valerio Miceli Barone, Fazl Barez, Shay B. Cohen, and Ioannis Konstas · 2023
Later among the works it cites.
Can large language models identify and reason about security vulnerabilities? Not yet, December 2023
Saad Ullah, Mingji Han, Saurabh Pujar, Hammond Pearce, Ayse Coskun, and Gianluca Stringhini · 2023
Later among the works it cites.
Data quality for software vulnerability datasets
Roland Croft, M Ali Babar, and Mehdi Kholoosi · 2023
Later among the works it cites.
Large language models for software engineering: A systematic literature review, 2024
Xinyi Hou, Yanjie Zhao, Yue Liu, Zhou Yang, Kailong Wang, Li Li, Xiapu Luo, David Lo, John Grundy, and Haoyu Wang · 2024
Closest in time.
Gpt-4 technical report, 2024
OpenAI · 2024
Closest in time.
Hugging face datasets, 2024
Benjamin Steenhoek · 2024
Closest in time.
Alex Gu, Wen-Ding Li, Naman Jain, Theo X Olausson, Celine Lee, Koushik Sen, and Armando Solar-Lezama · 2024
Closest in time.
Enhancing static analysis for practical bug detection: An llm-integrated approach
HAONAN LI, YU HAO, YIZHUO ZHAI, and ZHIYUN QIAN · 2024
Closest in time.
Llm4vuln: A unified evaluation framework for decoupling and enhancing llms’ vulnerability reasoning
Yuqiang Sun, Daoyuan Wu, Yue Xue, Han Liu, Wei Ma, Lyuye Zhang, Miaolei Shi, and Yang Liu · 2024
Closest in time.