Fetching the paper…
Reading the bibliography…
Large language models (large LMs) are increasingly trained on massive codebases and used to generate code.
CTRL: a Conditional Transformer Language Model for Controllable Generation
Nitish Shirish Keskar, Bryan McCann, Lav R. Varshney, Caiming Xiong, and Richard Socher. 2019 · 1909
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Questions Developers Ask While Diagnosing Potential Security Vulnerabilities with Static Analysis. In ESEC/FSE
Justin Smith, Brittany Johnson, Emerson R. Murphy-Hill, Bill Chu, and Heather Richter Lipford. 2015 · 2015
Earlier work this paper cites.
LAVA: Large-Scale Automated Vulnerability Addition. In IEEE S&P
Brendan Dolan-Gavitt, Patrick Hulin, Engin Kirda, Tim Leek, Andrea Mambretti, William K. Robertson, Frederick Ulrich, and Ryan Whelan. 2016 · 2016
Earlier work this paper cites.
Attention is All you Need. In NeurIPS
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Automatic Software Repair: a Survey. In ICSE
Luca Gazzola, Daniela Micucci, and Leonardo Mariani. 2018 · 2018
Earlier work this paper cites.
VulDeePecker: A Deep Learning-Based System for Vulnerability Detection. In NDSS
Zhen Li, Deqing Zou, Shouhuai Xu, Xinyu Ou, Hai Jin, Sujuan Wang, Zhijun Deng, and Yuyi Zhong. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
The Art, Science, and Engineering of Fuzzing: A Survey
Valentin J. M. Manès, HyungSeok Han, Choongwoo Han, Sang Kil Cha, Manuel Egele, Edward J. Schwartz, and Maverick Woo. 2021 · 2019
Earlier work this paper cites.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
Devign: Effective Vulnerability Identification by Learning Comprehensive Program Semantics via Graph Neural Networks. In NeurIPS
Yaqin Zhou, Shangqing Liu, Jing Kai Siow, Xiaoning Du, and Yang Liu. 2019 · 2019
Earlier work this paper cites.
Language Models are Few-Shot Learners. In NeurIPS
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Plug and Play Language Models: A Simple Approach to Controlled Text Generation. In ICLR
Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu. 2020 · 2020
Earlier work this paper cites.
A C/C++ Code Vulnerability Dataset with Code Changes and CVE Summaries. In MSR
Jiahao Fan, Yi Li, Shaohua Wang, and Tien N. Nguyen. 2020 · 2020
Earlier work this paper cites.
Software Vulnerability Detection Using Deep Neural Networks: A Survey
Guanjun Lin, Sheng Wen, Qing-Long Han, Jun Zhang, and Yang Xiang. 2020 · 2020
Earlier work this paper cites.
Program Synthesis with Large Language Models
Jacob Austin, Augustus Odena, Maxwell I. Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J. Cai, Michael Terry, Quoc V. Le, and Charles Sutton. 2021 · 2021
Earlier work this paper cites.
CVEfixes: Automated Collection of Vulnerabilities and Their Fixes from Open-source Software. In PROMISE
Guru Prasad Bhandari, Amara Naseer, and Leon Moonen. 2021 · 2021
Earlier work this paper cites.
Deep Learning Based Vulnerability Detection: Are We There Yet?
Saikat Chakraborty, Rahul Krishna, Yangruibo Ding, and Baishakhi Ray. 2022 · 2021
Earlier work this paper cites.
Evaluating Large Language Models Trained on Code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harrison Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Automatic Program Repair
Claire Le Goues, Michael Pradel, Abhik Roychoudhury, and Satish Chandra. 2021 · 2021
Earlier work this paper cites.
WARP: Word-level Adversarial ReProgramming. In ACL/IJCNLP
Karen Hambardzumyan, Hrant Khachatrian, and Jonathan May. 2021 · 2021
Earlier work this paper cites.
WILDS: A Benchmark of in-the-Wild Distribution Shifts. In ICML
Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, et al · 2021
Earlier work this paper cites.
GeDi: Generative Discriminator Guided Sequence Generation. In Findings of EMNLP
Ben Krause, Akhilesh Deepak Gotmare, Bryan McCann, Nitish Shirish Keskar, Shafiq R. Joty, Richard Socher, and Nazneen Fatema Rajani. 2021 · 2021
Earlier work this paper cites.
The Power of Scale for Parameter-Efficient Prompt Tuning. In EMNLP
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021 · 2021
Earlier work this paper cites.
Prefix-Tuning: Optimizing Continuous Prompts for Generation. In ACL/IJCNLP , Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (Eds.)
Xiang Lisa Li and Percy Liang. 2021 · 2021
Earlier work this paper cites.
SySeVR: A Framework for Using Deep Learning to Detect Software Vulnerabilities
Zhen Li, Deqing Zou, Shouhuai Xu, Hai Jin, Yawei Zhu, and Zhaoxuan Chen. 2022b · 2021
Earlier work this paper cites.
Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang. 2021 · 2021
Cited alongside, same era.
CrossVul: a Cross-language Vulnerability Dataset with Commit Data. In ESEC/FSE
Georgios Nikitopoulos, Konstantina Dritsa, Panos Louridas, and Dimitris Mitropoulos. 2021 · 2021
Cited alongside, same era.
Learning How to Ask: Querying LMs with Mixtures of Soft Prompts. In NAACL
Guanghui Qin and Jason Eisner. 2021 · 2021
Cited alongside, same era.
A Ground-truth Dataset of Real Security Patches
Sofia Reis and Rui Abreu. 2021 · 2021
Cited alongside, same era.
You Autocomplete Me: Poisoning Vulnerabilities in Neural Code Completion. In USENIX Security
Roei Schuster, Congzheng Song, Eran Tromer, and Vitaly Shmatikov. 2021 · 2021
Cited alongside, same era.
A Systematic Evaluation of Large Language Models of Code. In MAPS@PLDI
Frank F. Xu, Uri Alon, Graham Neubig, and Vincent Josua Hellendoorn. 2022 · 2022
Later among the works it cites.
FIXREVERTER: A Realistic Bug Injection Methodology for Benchmarking Fuzz Testing. In USENIX Security
Zenong Zhang, Zach Patterson, Michael Hicks, and Shiyi Wei. 2022 · 2022
Later among the works it cites.
AI Assistant for software developers | Tabnine
2023 · 2023
Closest in time.
AI Code Generator - Amazon CodeWhisperer - AWS
2023 · 2023
Closest in time.
fgets - cppreference.com
2023 · 2023
Closest in time.
Ghostwriter - Code faster with AI
2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation. In EMNLP
Yue Wang, Weishi Wang, Shafiq R. Joty, and Steven C. H. Hoi. 2021 · 2021
Cited alongside, same era.
VUDENC: Vulnerability Detection with Deep Learning on a Natural Codebase for Python
Laura Wartschinski, Yannic Noller, Thomas Vogel, Timo Kehrer, and Lars Grunske. 2022 · 2021
Cited alongside, same era.
2022 CWE Top 25 Most Dangerous Software Weaknesses
2022 · 2022
Cited alongside, same era.
Transcending TRANSCEND: Revisiting Malware Classification in the Presence of Concept Drift. In IEEE S&P
Federico Barbero, Feargus Pendlebury, Fabio Pierazzi, and Lorenzo Cavallaro. 2022 · 2022
Cited alongside, same era.
Efficient Training of Language Models to Fill in the Middle
Mohammad Bavarian, Heewoo Jun, Nikolas Tezak, John Schulman, Christine McLeavey, Jerry Tworek, and Mark Chen. 2022 · 2022
Cited alongside, same era.
MultiPL-E: A Scalable and Extensible Approach to Benchmarking Neural Code Generation
Federico Cassano, John Gouwar, Daniel Nguyen, Sydney Nguyen, Luna Phipps-Costin, Donald Pinckney, Ming-Ho Yee, Yangtian Zi, Carolyn Jane Anderson, Molly Q Feldman, Arjun Guha, Michael Greenberg, and Abhinav Jangda. 2022 · 2022
Cited alongside, same era.
Neural Transfer Learning for Repairing Security Vulnerabilities in C Code
Zimin Chen, Steve Kommrusch, and Martin Monperrus. 2023 · 2022
Cited alongside, same era.
2023 · 2023
Closest in time.
malloc - cppreference.com
2023 · 2023
Closest in time.
MarkupSafe · PyPI
2023 · 2023
Closest in time.
Models - Hugging Face
2023 · 2023
Closest in time.
PyYAML Documentation
2023 · 2023
Closest in time.
safe_join - Flask API
2023 · 2023
Closest in time.
The diff-match-patch Library
2023 · 2023
Closest in time.
Wikipedia - Common Weakness Enumeration
2023 · 2023
Closest in time.
Wikipedia - Kullback–Leibler Divergence
2023 · 2023
Closest in time.
SantaCoder: Don’t Reach for the Stars!
Loubna Ben Allal, Raymond Li, Denis Kocetkov, Chenghao Mou, Christopher Akiki, Carlos Muñoz Ferrandis, Niklas Muennighoff, Mayank Mishra, Alex Gu, Manan Dey, et al · 2023
Closest in time.
Data Quality for Software Vulnerability Datasets. In ICSE
Roland Croft, Muhammad Ali Babar, and M. Mehdi Kholoosi. 2023 · 2023
Closest in time.
GitHub Copilot X: the AI-powered Developer Experience
Thomas Dohmke. 2023 · 2023
Closest in time.
InCoder: A Generative Model for Code Infilling and Synthesis. In ICLR
Daniel Fried, Armen Aghajanyan, Jessy Lin, Sida Wang, Eric Wallace, Freda Shi, Ruiqi Zhong, Wen-tau Yih, Luke Zettlemoyer, and Mike Lewis. 2023 · 2023
Closest in time.
How Secure is Code Generated by ChatGPT?
Raphaël Khoury, Anderson R. Avila, Jacob Brunelle, and Baba Mamadou Camara. 2023 · 2023
Closest in time.
CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis. In ICLR
Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. 2023 · 2023
Closest in time.
Examining Zero-Shot Vulnerability Repair with Large Language Models. In IEEE S&P
Hammond Pearce, Benjamin Tan, Baleegh Ahmad, Ramesh Karri, and Brendan Dolan-Gavitt. 2023 · 2023
Closest in time.
Lost at C: A User Study on the Security Implications of Large Language Model Code Assistants. In USENIX Security
Gustavo Sandoval, Hammond Pearce, Teo Nys, Ramesh Karri, Siddharth Garg, and Brendan Dolan-Gavitt. 2023 · 2023
Closest in time.
StarCoder: May the source be with you!
John Smith. 2023 · 2023
Closest in time.
GitHub Copilot Now Has a Better AI Model and New Capabilities
Shuyin Zhao. 2023 · 2023
Closest in time.