Fetching the paper…
Reading the bibliography…
The integration of Large Language Models (LLMs) into software development has revolutionized the field, particularly through the use of Retrieval-Augmented Code Generation (RACG) systems that enhance code generation with information from external knowledge bases.
Some methods for classification and analysis of multivariate observations
James MacQueen et al · 1967
Earlier work this paper cites.
A statistical interpretation of term specificity and its application in retrieval
Karen Sparck Jones · 1972
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
The probabilistic relevance framework: Bm25 and beyond
Stephen Robertson, Hugo Zaragoza, et al · 2009
Earlier work this paper cites.
Poisoning attacks against support vector machines
Battista Biggio, Blaine Nelson, and Pavel Laskov · 2012
Earlier work this paper cites.
Juliet 1. 1 c/c++ and java test suite
Tim Boland and Paul E Black · 2012
Earlier work this paper cites.
Ebk-means: A clustering technique based on elbow method and k-means in wsn
Purnima Bholowalia and Arvind Kumar · 2014
Earlier work this paper cites.
Wild patterns: Ten years after the rise of adversarial machine learning
Battista Biggio and Fabio Roli · 2018
Earlier work this paper cites.
A software assurance reference dataset: Thousands of programs with known bugs
Paul E Black · 2018
Earlier work this paper cites.
Trojaning attack on neural networks
Yingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee, Juan Zhai, Weihang Wang, and Xiangyu Zhang · 2018
Earlier work this paper cites.
Poison frogs! targeted clean-label poisoning attacks on neural networks
Ali Shafahi, W Ronny Huang, Mahyar Najibi, Octavian Suciu, Christoph Studer, Tudor Dumitras, and Tom Goldstein · 2018
Earlier work this paper cites.
Integration k-means clustering method and elbow method for identification of the best customer profile cluster
Muhammad Ali Syakur, B Khusnul Khotimah, EMS Rochman, and Budi Dwi Satoto · 2018
Earlier work this paper cites.
Codesearchnet challenge: Evaluating the state of semantic code search
Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Introduction to the k-means clustering algorithm based on the elbow method
Mengyao Cui et al · 2020
Earlier work this paper cites.
Empirical review of automated analysis tools on 47,587 ethereum smart contracts
Thomas Durieux, João F Ferreira, Rui Abreu, and Pedro Cruz · 2020
Earlier work this paper cites.
Ac/c++ code vulnerability dataset with code changes and cve summaries
Jiahao Fan, Yi Li, Shaohua Wang, and Tien N Nguyen · 2020
Earlier work this paper cites.
How effective are smart contract analysis tools? evaluating smart contract static analysis tools using bug injection
Asem Ghaleb and Karthik Pattabiraman · 2020
Earlier work this paper cites.
Determine the number of unknown targets in open world based on elbow method
Fan Liu and Yong Deng · 2020
Earlier work this paper cites.
Generalizing from a few examples: A survey on few-shot learning
Yaqing Wang, Quanming Yao, James T Kwok, and Lionel M Ni · 2020
Earlier work this paper cites.
Cvefixes: automated collection of vulnerabilities and their fixes from open-source software
Guru Bhandari, Amara Naseer, and Leon Moonen · 2021
Earlier work this paper cites.
Deep learning based vulnerability detection: Are we there yet?
Saikat Chakraborty, Rahul Krishna, Yangruibo Ding, and Baishakhi Ray · 2021
Earlier work this paper cites.
Opportunities and challenges in code search tools
Chao Liu, Xin Xia, David Lo, Cuiyun Gao, Xiaohu Yang, and John Grundy · 2021
Earlier work this paper cites.
Retrieval augmented code generation and summarization
Md Rizwan Parvez, Wasi Uddin Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang · 2021
Earlier work this paper cites.
D2a: A dataset built for ai-based vulnerability detection methods using differential analysis
Yunhui Zheng, Saurabh Pujar, Burn Lewis, Luca Buratti, Edward Epstein, Bo Yang, Jim Laredo, Alessandro Morari, and Zhong Su · 2021
Earlier work this paper cites.
Crystalbleu: precisely and efficiently measuring the similarity of code
Aryaz Eghbali and Michael Pradel · 2022
Earlier work this paper cites.
Re2g: Retrieve, rerank, generate
Michael Glass, Gaetano Rossiello, Md Faisal Mahbub Chowdhury, Ankita Rajaram Naik, Pengshan Cai, and Alfio Gliozzo · 2022
Earlier work this paper cites.
Jigsaw: Large language models meet program synthesis
Naman Jain, Skanda Vaidyanath, Arun Iyer, Nagarajan Natarajan, Suresh Parthasarathy, Sriram Rajamani, and Rahul Sharma · 2022
Earlier work this paper cites.
Asleep at the keyboard? assessing the security of github copilot’s code contributions
Hammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt, and Ramesh Karri · 2022
Earlier work this paper cites.
Ignore previous prompt: Attack techniques for language models
Fábio Perez and Ian Ribeiro · 2022
Earlier work this paper cites.
Diversevul: A new vulnerable source code dataset for deep learning based vulnerability detection
Yizheng Chen, Zhoujie Ding, Lamya Alowain, Xinyun Chen, and David Wagner · 2023
Earlier work this paper cites.
Retrieval-augmented generation for large language models: A survey
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, and Haofen Wang · 2023
Cited alongside, same era.
How secure is code generated by chatgpt?
Raphaël Khoury, Anderson R Avila, Jacob Brunelle, and Baba Mamadou Camara · 2023
Cited alongside, same era.
Multi-step jailbreaking privacy attacks on chatgpt
Haoran Li, Dadi Guo, Wei Fan, Mingshi Xu, Jie Huang, Fanpu Meng, and Yangqiu Song · 2023
Cited alongside, same era.
From classification to generation: Insights into crosslingual retrieval augmented icl
Xiaoqian Li, Ercong Nie, and Sheng Liang · 2023
Cited alongside, same era.
Prompt injection attacks and defenses in llm-integrated applications
Preference-guided refactored tuning for retrieval augmented code generation
Xinyu Gao, Yun Xiong, Deze Wang, Zhenhan Guan, Zejian Shi, Haofen Wang, and Shanshan Li · 2024
Later among the works it cites.
Large language models are few-shot summarizers: Multi-intent comment generation via in-context learning
Mingyang Geng, Shangwen Wang, Dezun Dong, Haotian Wang, Ge Li, Zhi Jin, Xiaoguang Mao, and Xiangke Liao · 2024
Later among the works it cites.
Hui Huang, Yingqi Qu, Jing Liu, Muyun Yang, and Tiejun Zhao · 2024
Later among the works it cites.
Using ai assistants in software development: A qualitative study on security practices and concerns
Jan H Klemmer, Stefan Albert Horstmann, Nikhil Patnaik, Cordelia Ludden, Cordell Burton Jr, Carson Powers, Fabio Massacci, Akond Rahman, Daniel Votipka, Heather Richter Lipford, et al · 2024
Later among the works it cites.
An empirical study on low code programming using traditional vs large language model support
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia, and Neil Zhenqiang Gong · 2023
Cited alongside, same era.
Query rewriting for retrieval-augmented large language models
Xinbei Ma, Yeyun Gong, Pengcheng He, Hai Zhao, and Nan Duan · 2023
Cited alongside, same era.
National vulnerability database (nvd)
National Institute of Standards and Technology (NIST) · 2023
Cited alongside, same era.
Rodrigo Pedro, Daniel Castro, Paulo Carreira, and Nuno Santos · 2023
Cited alongside, same era.
Evaluating and optimizing the effectiveness of neural machine translation in supporting code retrieval models: A study on the cat benchmark
Hung Phan and Ali Jannesari · 2023
Cited alongside, same era.
Code llama: Open foundation models for code
Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Romain Sauvestre, Tal Remez, et al · 2023
Cited alongside, same era.
A comprehensive survey of few-shot learning: Evolution, applications, challenges, and opportunities
Yisheng Song, Ting Wang, Puyu Cai, Subrota K Mondal, and Jyoti Prakash Sahoo · 2023
Cited alongside, same era.
Efficient avoidance of vulnerabilities in auto-completed smart contract code using vulnerability-constrained decoding
André Storhaug, Jingyue Li, and Tianyuan Hu · 2023
Cited alongside, same era.
Yongkun Liu, Jiachi Chen, Tingting Bi, John Grundy, Yanlin Wang, Jianxing Yu, Ting Chen, Yutian Tang, and Zibin Zheng · 2024
Later among the works it cites.
Formalizing and benchmarking prompt injection attacks and defenses
Yupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia, and Neil Zhenqiang Gong · 2024
Later among the works it cites.
Github language statistics (githut), 2024
madnight · 2024
Later among the works it cites.
Kragen: a knowledge graph-enhanced rag framework for biomedical problem solving using large language models
Nicholas Matsumoto, Jay Moran, Hyunjun Choi, Miguel E Hernandez, Mythreye Venkatesan, Paul Wang, and Jason H Moore · 2024
Later among the works it cites.
Retrieval-augmented generation (rag) in azure machine learning, 2024
Microsoft · 2024
Later among the works it cites.
2024 cwe top 25 most dangerous software weaknesses
MITRE Corporation · 2024
Later among the works it cites.
Ollama framework
Ollama · 2024
Later among the works it cites.
Chatgpt retrieval plugin, 2024
OpenAI · 2024
Later among the works it cites.
Openai api reference - chat create n, 2024
OpenAI · 2024
Later among the works it cites.
Openai models - embeddings
OpenAI · 2024
Later among the works it cites.
An empirical study of the non-determinism of chatgpt in code generation
Shuyin Ouyang, Jie M Zhang, Mark Harman, and Meng Wang · 2024
Later among the works it cites.
Visual adversarial examples jailbreak aligned large language models
Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Peter Henderson, Mengdi Wang, and Prateek Mittal · 2024
Later among the works it cites.
Rag-fusion: a new take on retrieval-augmented generation
Zackary Rackauckas · 2024
Later among the works it cites.
Llm safety leaderboard
Secure Learning Lab · 2024
Later among the works it cites.
A systematic literature review on automated software vulnerability detection using machine learning
Nima Shiri Harzevili, Alvine Boaye Belle, Junjie Wang, Song Wang, Zhen Ming Jiang, and Nachiappan Nagappan · 2024
Later among the works it cites.
jina-embeddings-v3: Multilingual embeddings with task lora
Saba Sturua, Isabelle Mohr, Mohammad Kalim Akram, Michael Günther, Bo Wang, Markus Krimmel, Feng Wang, Georgios Mastrapas, Andreas Koukounas, Nan Wang, et al · 2024
Later among the works it cites.
Evor: Evolving retrieval for code generation
Hongjin Su, Shuyang Jiang, Yuhang Lai, Haoyuan Wu, Boao Shi, Che Liu, Qian Liu, and Tao Yu · 2024
Later among the works it cites.
Small models, big insights: Leveraging slim proxy models to decide when and what to retrieve for LLMs
Jiejun Tan, Zhicheng Dou, Yutao Zhu, Peidong Guo, Kun Fang, and Ji-Rong Wen · 2024
Later among the works it cites.
Fusing code searchers
Shangwen Wang, Mingyang Geng, Bo Lin, Zhensu Sun, Ming Wen, Yepang Liu, Li Li, Tegawendé F Bissyandé, and Xiaoguang Mao · 2024
Later among the works it cites.
Benchmark self-evolving: A multi-agent framework for dynamic llm evaluation
Siyuan Wang, Zhuohan Long, Zhihao Fan, Zhongyu Wei, and Xuanjing Huang · 2024
Later among the works it cites.
Reposvul: A repository-level high-quality vulnerability dataset
Xinchen Wang, Ruida Hu, Cuiyun Gao, Xin-Cheng Wen, Yujia Chen, and Qing Liao · 2024
Later among the works it cites.
Coderag-bench: Can retrieval augment code generation?
Zora Zhiruo Wang, Akari Asai, Xinyan Velocity Yu, Frank F Xu, Yiqing Xie, Graham Neubig, and Daniel Fried · 2024
Later among the works it cites.
Hijackrag: Hijacking attacks against retrieval-augmented large language models
Yucheng Zhang, Qinfeng Li, Tianyu Du, Xuhong Zhang, Xinkui Zhao, Zhengwen Feng, and Jianwei Yin · 2024
Later among the works it cites.
Deepseek-coder-v2: Breaking the barrier of closed-source models in code intelligence
Qihao Zhu, Daya Guo, Zhihong Shao, Dejian Yang, Peiyi Wang, Runxin Xu, Y Wu, Yukun Li, Huazuo Gao, Shirong Ma, et al · 2024
Later among the works it cites.
Poisonedrag: Knowledge poisoning attacks to retrieval-augmented generation of large language models
Wei Zou, Runpeng Geng, Binghui Wang, and Jinyuan Jia · 2024
Later among the works it cites.
How secure is ai-generated code: a large-scale comparison of large language models
Norbert Tihanyi, Tamas Bisztray, Mohamed Amine Ferrag, Ridhi Jain, and Lucas C Cordeiro · 2025
Closest in time.