Fetching the paper…
Reading the bibliography…
The significant advancements in Large Language Models (LLMs) have resulted in their widespread adoption across various tasks within Software Engineering (SE), including vulnerability detection and repair.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 1901
Earlier work this paper cites.
Identifying relevant studies in software engineering
Zhang, H., Babar, M. A., and Tell, P · 2011
Earlier work this paper cites.
Mitigating program security vulnerabilities: Approaches and challenges
Shahriar, H., and Zulkernine, M · 2012
Earlier work this paper cites.
Guidelines for snowballing in systematic literature studies and a replication in software engineering
Wohlin, C · 2014
Earlier work this paper cites.
Software vulnerability analysis and discovery using machine-learning and data-mining techniques: A survey
Ghaffarian, S. M., and Shahriari, H. R · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Learning and evaluating contextual embedding of source code
Kanade, A., Maniatis, P., Balakrishnan, G., and Shi, K · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Codebert: A pre-trained model for programming and natural languages
Feng, Z., Guo, D., Tang, D., Duan, N., Feng, X., Gong, M., Shou, L., Qin, B., Liu, T., Jiang, D., et al · 2020
Earlier work this paper cites.
Graphcodebert: Pre-training code representations with data flow
Guo, D., Ren, S., Lu, S., Feng, Z., Tang, D., Liu, S., Zhou, L., Duan, N., Svyatkovskiy, A., Fu, S., et al · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive NLP tasks
Lewis, P. S. H., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., and Kiela, D · 2020
Earlier work this paper cites.
Software vulnerability detection using deep neural networks: a survey
Lin, G., Wen, S., Han, Q.-L., Zhang, J., and Xiang, Y · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Codebleu: a method for automatic evaluation of code synthesis
Ren, S., Guo, D., Lu, S., Zhou, L., Liu, S., Tang, D., Sundaresan, N., Zhou, M., Blanco, A., and Ma, S · 2020
Earlier work this paper cites.
Code and named entity recognition in stackoverflow
Tabassum, J., Maddela, M., Xu, W., and Ritter, A · 2020
Earlier work this paper cites.
Bertscore: Evaluating text generation with BERT
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y · 2020
Earlier work this paper cites.
Unified pre-training for program understanding and generation
Ahmad, W. U., Chakraborty, S., Ray, B., and Chang, K.-W · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Earlier work this paper cites.
Codexglue: A machine learning benchmark dataset for code understanding and generation
Lu, S., Guo, D., Ren, S., Huang, J., Svyatkovskiy, A., Blanco, A., Clement, C., Drain, D., Jiang, D., Tang, D., et al · 2021
Earlier work this paper cites.
Wang, Y., Wang, W., Joty, S., and Hoi, S. C · 2021
Earlier work this paper cites.
Security vulnerability detection using deep learning natural language processing
Ziems, N., and Wu, S · 2021
Earlier work this paper cites.
Natgen: generative pre-training by "naturalizing" source code
Chakraborty, S., Ahmed, T., Ding, Y., Devanbu, P. T., and Ray, B · 2022
Earlier work this paper cites.
Linevul: A transformer-based line-level vulnerability prediction
Fu, M., and Tantithamthavorn, C · 2022
Earlier work this paper cites.
Vulrepair: a t5-based automated software vulnerability repair
Fu, M., Tantithamthavorn, C., Le, T., Nguyen, V., and Phung, D. Q · 2022
Earlier work this paper cites.
Unixcoder: Unified cross-modal pre-training for code representation
Guo, D., Lu, S., Duan, N., Wang, Y., Zhou, M., and Yin, J · 2022
Earlier work this paper cites.
Vulberta: Simplified source code pre-training for vulnerability detection
Hanif, H., and Maffeis, S · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2022
Earlier work this paper cites.
We analysed 90,000+ software vulnerabilities: Here’s what we learned
TARGETT, E · 2022
Earlier work this paper cites.
Transformer-based language models for software vulnerability detection
Thapa, C., Jang, S. I., Ahmed, M. E., Camtepe, S., Pieprzyk, J., and Nepal, S · 2022
Earlier work this paper cites.
Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions
Trivedi, H., Balasubramanian, N., Khot, T., and Sabharwal, A · 2022
Earlier work this paper cites.
Code vulnerability detection based on deep sequence and graph models: A survey
Wu, B., Zou, F., et al · 2022
Earlier work this paper cites.
A systematic evaluation of large language models of code
Xu, F. F., Alon, U., Neubig, G., and Hellendoorn, V. J · 2022
Earlier work this paper cites.
Natural attack for pre-trained models of code
Yang, Z., Shi, J., He, J., and Lo, D · 2022
Earlier work this paper cites.
Spvf: security property assisted vulnerability fixing via attention-based models
Zhou, Z., Bo, L., Wu, X., Sun, X., Zhang, T., Li, B., Zhang, J., and Cao, S · 2022
Earlier work this paper cites.
Fixing hardware security bugs with large language models
Ahmad, B., Thakur, S., Tan, B., Karri, R., and Pearce, H · 2023
Earlier work this paper cites.
Self-rag: Learning to retrieve, generate, and critique through self-reflection
Asai, A., Wu, Z., Wang, Y., Sil, A., and Hajishirzi, H · 2023
Earlier work this paper cites.
Diversevul: A new vulnerable source code dataset for deep learning based vulnerability detection
Chen, Y., Ding, Z., Alowain, L., Chen, X., and Wagner, D. A · 2023
Earlier work this paper cites.
Neural transfer learning for repairing security vulnerabilities in C code
Chen, Z., Kommrusch, S., and Monperrus, M · 2023
Earlier work this paper cites.
Seqtrans: Automatic vulnerability fix via sequence to sequence learning
Chi, J., Qu, Y., Liu, T., Zheng, Q., and Yin, H · 2023
Earlier work this paper cites.
Data quality for software vulnerability datasets
Croft, R., Babar, M. A., and Kholoosi, M. M · 2023
Cited alongside, same era.
Incoder: A generative model for code infilling and synthesis
Fried, D., Aghajanyan, A., Lin, J., Wang, S., Wallace, E., Shi, F., Zhong, R., Yih, S., Zettlemoyer, L., and Lewis, M · 2023
Cited alongside, same era.
Chatgpt for vulnerability detection, classification, and repair: How far are we?
Fu, M., Tantithamthavorn, C., Nguyen, V., and Le, T · 2023
Cited alongside, same era.
Github copilot
GitHub · 2023
Cited alongside, same era.
Large language models for code: Security hardening and adversarial testing
He, J., and Vechev, M · 2023
Cited alongside, same era.
Representation learning for stack overflow posts: How far are we?
He, J., Zhou, X., Xu, B., Zhang, T., Kim, K., Yang, Z., Thung, F., Irsan, I. C., and Lo, D · 2023
Cited alongside, same era.
Unifying the perspectives of nlp and software engineering: A survey on language models for code
Zhang, Z., Chen, C., Liu, B., Liao, C., Gong, Z., Yu, H., Li, J., and Wang, R · 2023
Later among the works it cites.
A survey of large language models
Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., et al · 2023
Later among the works it cites.
The devil is in the tails: How long-tailed code distributions impact large language models
Zhou, X., Kim, K., Xu, B., Liu, J., Han, D., and Lo, D · 2023
Later among the works it cites.
Ccbert: Self-supervised code change representation learning
Zhou, X., Xu, B., Han, D., Yang, Z., He, J., and Lo, D · 2023
Later among the works it cites.
https://docs.google.com/document/d/18-UrkfH35CNMGRjjsDYZGK6L1aC9wP3GsKCtrIekcUQ/edit?usp=sharing , 2024
Online Appendix for This Review · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Large language models for software engineering: A systematic literature review
Hou, X., Zhao, Y., Liu, Y., Yang, Z., Wang, K., Li, L., Luo, X., Lo, D., Grundy, J. C., and Wang, H · 2023
Cited alongside, same era.
Understanding the effectiveness of large language models in detecting security vulnerabilities
Khare, A., Dutta, S., Li, Z., Solko-Breslin, A., Alur, R., and Naik, M · 2023
Cited alongside, same era.
Leveraging user-defined identifiers for counterfactual data generation in source code vulnerability detection
Kuang, H., Yang, F., Zhang, L., Tang, G., and Yang, L · 2023
Cited alongside, same era.
Comparison and evaluation on static application security testing (sast) tools for java
Li, K., Chen, S., Fan, L., Feng, R., Liu, H., Liu, C., Liu, Y., and Chen, Y · 2023
Cited alongside, same era.
Software vulnerability detection with GPT and in-context learning
Liu, Z., Liao, Q., Gu, W., and Gao, C · 2023
Cited alongside, same era.
Trustworthy and synergistic artificial intelligence for software engineering: Vision and roadmaps
Lo, D · 2023
Cited alongside, same era.
Closest in time.
Deepcode AI fix: Fixing security vulnerabilities with large language models
Berabi, B., Gronskiy, A., Raychev, V., Sivanrupan, G., Chibotaru, V., and Vechev, M. T · 2024
Closest in time.
Vulnerability detection with code language models: How far are we?
Ding, Y., Fu, Y., Ibrahim, O., Sitawarin, C., Chen, X., Alomair, B., Wagner, D. A., Ray, B., and Chen, Y · 2024
Closest in time.
Vul-rag: Enhancing llm-based vulnerability detection via knowledge-level RAG
Du, X., Zheng, G., Wang, K., Feng, J., Deng, W., Liu, M., Chen, B., Peng, X., Ma, T., and Lou, Y · 2024
Closest in time.
Vision transformer-inspired automated vulnerability repair
Fu, M., Nguyen, V., Tantithamthavorn, C., Phung, D., and Le, T · 2024
Closest in time.
Aibughunter: A practical tool for predicting, classifying and repairing software vulnerabilities
Fu, M., Tantithamthavorn, C., Le, T., Kume, Y., Nguyen, V., Phung, D. Q., and Grundy, J. C · 2024
Closest in time.
Llm-powered code vulnerability repair with reinforcement learning and semantic reward
Islam, N. T., Khoury, J., Seong, A., Parra, G. D. L. T., Bou-Harb, E., and Najafirad, P · 2024
Closest in time.
Code security vulnerability repair using reinforcement learning with large language models
Islam, N. T., and Najafirad, P · 2024
Closest in time.
Investigating data contamination for pre-training language models, 2024
Jiang, M., Liu, K. Z., Zhong, M., Schaeffer, R., Ouyang, S., Han, J., and Koyejo, S · 2024
Closest in time.
DFEPT: data flow embedding for enhancing pre-trained model based vulnerability detection
Jiang, Z., Sun, W., Gu, X., Wu, J., Wen, T., Hu, H., and Yan, M · 2024
Closest in time.
Are latent vulnerabilities hidden gems for software vulnerability prediction? an empirical study
Le, T. H. M., Du, X., and Babar, M. A · 2024
Closest in time.
On the effectiveness of function-level vulnerability detectors for inter-procedural vulnerabilities
Li, Z., Wang, N., Zou, D., Li, Y., Zhang, R., Xu, S., Zhang, C., and Jin, H · 2024
Closest in time.
On the effectiveness of function-level vulnerability detectors for inter-procedural vulnerabilities
Li, Z., Wang, N., Zou, D., Li, Y., Zhang, R., Xu, S., Zhang, C., and Jin, H · 2024
Closest in time.
A large-scale survey on the usability of AI programming assistants: Successes and challenges
Liang, J. T., Yang, C., and Myers, B. A · 2024
Closest in time.
Pre-training by predicting program dependencies for vulnerability analysis tasks
Liu, Z., Tang, Z., Zhang, J., Xia, X., and Yang, X · 2024
Closest in time.
Towards causal deep learning for vulnerability detection
Mahbubur Rahman, M., Ceka, I., Mao, C., Chakraborty, S., Ray, B., and Le, W · 2024
Closest in time.
Microsoft copilot for security
Microsoft · 2024
Closest in time.
Large language models: A survey, 2024
Minaee, S., Mikolov, T., Nikzad, N., Chenaghlu, M., Socher, R., Amatriain, X., and Gao, J · 2024
Closest in time.
Learning-based models for vulnerability detection: An extensive study
Ni, C., Shen, L., Xu, X., Yin, X., and Wang, S · 2024
Closest in time.
Nong, Y., Aldeen, M., Cheng, L., Hu, H., Chen, F., and Cai, H · 2024
Closest in time.
Automated software vulnerability patching using large language models
Nong, Y., Yang, H., Cheng, L., Hu, H., and Cai, H · 2024
Closest in time.
Uncovering the limits of machine learning for automatic vulnerability detection
Risse, N., and Böhme, M · 2024
Closest in time.
Toward improved deep learning-based vulnerability detection
Sejfia, A., Das, S., Shafiq, S., and Medvidovic, N · 2024
Closest in time.
Toward improved deep learning-based vulnerability detection
Sejfia, A., Das, S., Shafiq, S., and Medvidović, N · 2024
Closest in time.
Finetuning large language models for vulnerability detection
Shestov, A., Cheshkov, A., Levichev, R., Mussabayev, R., Zadorozhny, P., Maslov, E., Vadim, C., and Bulychev, E · 2024
Closest in time.
Detectvul: A statement-level code vulnerability detection for python
Tran, H.-C., Tran, A.-D., and Le, K.-H · 2024
Closest in time.
Combining structured static code information and dynamic symbolic traces for software vulnerability prediction
Wang, H., Tang, Z., Tan, S. H., Wang, J., Liu, Y., Fang, H., Xia, C., and Wang, Z · 2024
Closest in time.
A survey on large language model based autonomous agents
Wang, L., Ma, C., Feng, X., Zhang, Z., Yang, H., Zhang, J., Chen, Z., Tang, J., Chen, X., Lin, Y., et al · 2024
Closest in time.
Navrepair: Node-type aware C/C++ code vulnerability repair
Wang, R., Li, Z., Wang, C., Xiao, Y., and Gao, C · 2024
Closest in time.
Vuleval: Towards repository-level evaluation of software vulnerability detection
Wen, X., Wang, X., Chen, Y., Hu, R., Lo, D., and Gao, C · 2024
Closest in time.
Matsvd: Boosting statement-level vulnerability detection via dependency-based attention
Weng, C., Qin, Y., Lin, B., Liu, P., and Chen, L · 2024
Closest in time.
Security vulnerability detection with multitask self-instructed fine-tuning of large language models
Yang, A. Z., Tian, H., Ye, H., Martins, R., and Goues, C. L · 2024
Closest in time.
Harnessing the power of llms in practice: A survey on chatgpt and beyond
Yang, J., Jin, H., Tang, R., Han, X., Feng, Q., Jiang, H., Zhong, S., Yin, B., and Hu, X. B · 2024
Closest in time.
Pros and cons! evaluating chatgpt on software vulnerability
Yin, X · 2024
Closest in time.
Out of sight, out of mind: Better automatic vulnerability repair by broadening input ranges and sources
Zhou, X., Kim, K., Xu, B., Han, D., and Lo, D · 2024
Closest in time.
Zhou, X., Tran, D.-M., Le-Cong, T., Zhang, T., Irsan, I. C., Sumarlin, J., Le, B., and Lo, D · 2024
Closest in time.
Large language model for vulnerability detection: Emerging results and future directions
Zhou, X., Zhang, T., and Lo, D · 2024
Closest in time.