Fetching the paper…
Reading the bibliography…
This paper presents CyberSecEval, a comprehensive benchmark developed to help bolster the cybersecurity of Large Language Models (LLMs) employed as coding assistants.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Asleep at the keyboard? assessing the security of github copilot’s code contributions
Hammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt, and Ramesh Karri · 2022
Earlier work this paper cites.
Securityeval dataset: mining vulnerability examples to evaluate machine learning-based code generation techniques
Mohammed Latif Siddiq and Joanna CS Santos · 2022
Earlier work this paper cites.
An empirical study of code smells in transformer-based code generation techniques
Mohammed Latif Siddiq, Shafayat H. Majumder, Maisha R. Mim, Sourov Jajodia, and Joanna C. S. Santos · 2022
Earlier work this paper cites.
Github copilot for business is now available, Feb 2023
Thomas Dohmke · 2023
Earlier work this paper cites.
Systematically finding security vulnerabilities in black-box code generation models
Hossein Hajipour, Thorsten Holz, Lea Schönherr, and Mario Fritz · 2023
Cited alongside, same era.
How secure is code generated by chatgpt?
Raphaël Khoury, Anderson R Avila, Jacob Brunelle, and Baba Mamadou Camara · 2023
Cited alongside, same era.
Starcoder: may the source be with you!
Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, et al · 2023
Cited alongside, same era.
Common weakness enumeration: A community-developed list of software & hardware weakness types
Corporation MITRE · 2023
Cited alongside, same era.
Mitre att&ck®
Corporation MITRE · 2023
Cited alongside, same era.
Code llama: Open foundation models for code
Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, et al · 2023
Closest in time.
Lost at c: A user study on the security implications of large language model code assistants
Gustavo Sandoval, Hammond Pearce, Teo Nys, Ramesh Karri, Siddharth Garg, and Brendan Dolan-Gavitt · 2023
Closest in time.
The formai dataset: Generative ai in software security through the lens of formal verification
Norbert Tihanyi, Tamas Bisztray, Ridhi Jain, Mohamed Amine Ferrag, Lucas C Cordeiro, and Vasileios Mavroeidis · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Codecompose: A large-scale industrial deployment of ai-assisted code authoring
Vijayaraghavan Murali, Chandra Maddila, Imad Ahmad, Michael Bolin, Daniel Cheng, Negar Ghorbani, Renuka Fernandez, and Nachiappan Nagappan · 2023
Cited alongside, same era.
Burak Yetiştiren, Işık Özsoy, Miray Ayerdem, and Eray Tüzün · 2023
Closest in time.
A study on robustness and reliability of large language model code generation
Li Zhong and Zilong Wang · 2023
Closest in time.