Fetching the paper…
Reading the bibliography…
Semantic code search, retrieving code that matches a given natural language query, is an important task to improve productivity in software engineering.
Krippendorff, klaus, content analysis: An introduction to its methodology. beverly hills, ca: Sage, 1980
Klaus Krippendorff · 1980
Earlier work this paper cites.
A comparative analysis of functional correctness
Douglas D Dunlop and Victor R Basili · 1982
Earlier work this paper cites.
Code complete
Steve McConnell · 2004
Earlier work this paper cites.
Sourcerer: a search engine for open source code supporting structure-based search
Sushil Bajracharya, Trung Ngo, Erik Linstead, Yimeng Dou, Paul Rigor, Pierre Baldi, and Cristina Lopes · 2006
Earlier work this paper cites.
Clean code: a handbook of agile software craftsmanship
Robert C Martin · 2009
Earlier work this paper cites.
Understanding bag-of-words model: a statistical framework
Yin Zhang, Rong Jin, and Zhi-Hua Zhou · 2010
Earlier work this paper cites.
How well do search engines support code retrieval on the web?
Susan Elliott Sim, Medha Umarji, Sukanya Ratanotayanon, and Cristina V Lopes · 2011
Earlier work this paper cites.
What help do developers seek, when and how?
Hongwei Li, Zhenchang Xing, Xin Peng, and Wenyun Zhao · 2013
Earlier work this paper cites.
How to effectively use topic models for software engineering tasks? an approach based on genetic algorithms
Annibale Panichella, Bogdan Dit, Rocco Oliveto, Massimilano Di Penta, Denys Poshynanyk, and Andrea De Lucia · 2013
Earlier work this paper cites.
A theoretical analysis of ndcg type ranking measures
Yining Wang, Liwei Wang, Yuanzhi Li, Di He, and Tie-Yan Liu · 2013
Earlier work this paper cites.
Solving the search for source code
Kathryn T Stolee, Sebastian Elbaum, and Daniel Dobos · 2014
Earlier work this paper cites.
Staqc: A systematically mined question-code dataset from stack overflow
Ziyu Yao, Daniel S Weld, Wei-Peng Chen, and Huan Sun · 2018
Earlier work this paper cites.
Learning to mine aligned code and natural language pairs from stack overflow
Pengcheng Yin, Bowen Deng, Edgar Chen, Bogdan Vasilescu, and Graham Neubig · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Codesearchnet challenge: Evaluating the state of semantic code search
Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt · 2019
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le · 2019
Earlier work this paper cites.
Automatic query reformulation for code search using crowdsourced knowledge
Mohammad M Rahman, Chanchal K Roy, and David Lo · 2019
Earlier work this paper cites.
Are the code snippets what we are searching for? a benchmark and an empirical study on code search with natural-language queries
Shuhan Yan, Hang Yu, Yuting Chen, Beijun Shen, and Lingxiao Jiang · 2020
Earlier work this paper cites.
Codebert: A pre-trained model for programming and natural languages
Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, et al · 2020
Earlier work this paper cites.
PyMT5: multi-mode translation of natural language and python code with transformers
Colin Clement, Dawn Drain, Jonathan Timcheck, Alexey Svyatkovskiy, and Neel Sundaresan · 2020
Earlier work this paper cites.
Neural code search revisited: Enhancing code snippet retrieval through natural language intent
Geert Heyman and Tom Van Cutsem · 2020
Earlier work this paper cites.
Graphcodebert: Pre-training code representations with data flow
Daya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng, Duyu Tang, Shujie Liu, Long Zhou, Nan Duan, Alexey Svyatkovskiy, Shengyu Fu, et al · 2020
Earlier work this paper cites.
Electra: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Earlier work this paper cites.
Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers, 2020
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou · 2020
Cited alongside, same era.
Mpnet: Masked and permuted pre-training for language understanding
Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu · 2020
Cited alongside, same era.
Learning code-query interaction for enhancing code searches
Wei Li, Haozhe Qin, Shuhan Yan, Beijun Shen, and Yuting Chen · 2020
Cited alongside, same era.
Trans3̂: A transformer-based framework for unifying code summarization and code search, 2020
Wenhua Wang, Yuqun Zhang, Zhengran Zeng, and Guandong Xu · 2020
Cited alongside, same era.
Cosqa: 20,000+ web queries for code search and question answering
Junjie Huang, Duyu Tang, Linjun Shou, Ming Gong, Ke Xu, Daxin Jiang, Ming Zhou, and Nan Duan · 2021
Magicoder: Empowering code generation with oss-instruct
Yuxiang Wei, Zhe Wang, Jiawei Liu, Yifeng Ding, and Lingming Zhang · 2023
Later among the works it cites.
Retrieval-based prompt selection for code-related few-shot learning
Noor Nashid, Mifta Sintaha, and Ali Mesbah · 2023
Later among the works it cites.
Keep the conversation going: Fixing 162 out of 337 bugs for $0.42 each using chatgpt
Chunqiu Steven Xia and Lingming Zhang · 2023
Later among the works it cites.
Impact of code language models on automated program repair
Nan Jiang, Kevin Liu, Thibaud Lutellier, and Lin Tan · 2023
Later among the works it cites.
Automated program repair in the era of large pre-trained language models
Chunqiu Steven Xia, Yuxiang Wei, and Lingming Zhang · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Codexglue: A machine learning benchmark dataset for code understanding and generation
Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin Clement, Dawn Drain, Daxin Jiang, Duyu Tang, et al · 2021
Cited alongside, same era.
Codesc: A large code-description parallel dataset, 2021
Masum Hasan, Tanveer Muttaqueen, Abdullah Al Ishtiaq, Kazi Sajeed Mehrab, Md. Mahim Anjum Haque, Tahmid Hasan, Wasi Uddin Ahmad, Anindya Iqbal, and Rifat Shahriyar · 2021
Cited alongside, same era.
Beir: A heterogenous benchmark for zero-shot evaluation of information retrieval models
Nandan Thakur, Nils Reimers, Andreas Ruckl’e, Abhishek Srivastava, and Iryna Gurevych · 2021
Cited alongside, same era.
Yue Wang, Weishi Wang, Shafiq Joty, and Steven CH Hoi · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
J. Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen · 2021
Cited alongside, same era.
All-minilm-l12-v2
Nils Reimers and Iryna Gurevych · 2021
Cited alongside, same era.
All-mpnet-base-v2
Kexin Song et al · 2021
Cited alongside, same era.
Scoping software engineering for ai: the tse perspective
Sebastian Uchitel, Marsha Chechik, Massimiliano Di Penta, Bram Adams, Nazareno Aguirre, Gabriele Bavota, Domenico Bianculli, Kelly Blincoe, Ana Cavalcanti, Yvonne Dittrich, et al · 2024
Closest in time.
Optimizing code retrieval: High-quality and scalable dataset annotation through large language models
Rui Li, Qi Liu, Liyang He, Zheng Zhang, Hao Zhang, Shengyu Ye, Junyu Lu, and Zhenya Huang · 2024
Closest in time.
Procqa: A large-scale community-based programming question answering dataset for code search
Zehan Li, Jianfei Zhang, Chuantao Yin, Yuanxin Ouyang, and Wenge Rong · 2024
Closest in time.
Coir: A comprehensive benchmark for code information retrieval models
Xiangyang Li, Kuicai Dong, Yi Lee, Wei Xia, Yichun Yin, Hao Zhang, Yong Liu, Yasheng Wang, and Ruiming Tang · 2024
Closest in time.
jina-embeddings-v3: Multilingual embeddings with task lora, 2024
Saba Sturua, Isabelle Mohr, Mohammad Kalim Akram, Michael Günther, Bo Wang, Markus Krimmel, Feng Wang, Georgios Mastrapas, Andreas Koukounas, Nan Wang, and Han Xiao · 2024
Closest in time.
Multilingual e5 text embeddings: A technical report
Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei · 2024
Closest in time.
Large language models are edge-case generators: Crafting unusual programs for fuzzing deep learning libraries
Yinlin Deng, Chunqiu Steven Xia, Chenyuan Yang, Shizhuo Dylan Zhang, Shujing Yang, and Lingming Zhang · 2024
Closest in time.
Generative ai to generate test data generators
Benoit Baudry, Khashayar Etemadi, Sen Fang, Yogya Gamage, Yi Liu, Yuxin Liu, Martin Monperrus, Javier Ron, André Silva, and Deepika Tiwari · 2024
Closest in time.
Leveraging large language models for enhancing the understandability of generated unit tests
Amirhossein Deljouyi, Roham Koohestani, Maliheh Izadi, and Andy Zaidman · 2024
Closest in time.
Automated program repair, what is it good for? not absolutely nothing!
Hadeel Eladawy, Claire Le Goues, and Yuriy Brun · 2024
Closest in time.
Cigar: Cost-efficient program repair with llms
Dávid Hidvégi, Khashayar Etemadi, Sofia Bobadilla, and Martin Monperrus · 2024
Closest in time.
Mathematical discoveries from program search with large language models
Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Matej Balog, M Pawan Kumar, Emilien Dupont, Francisco JR Ruiz, Jordan S Ellenberg, Pengming Wang, Omar Fawzi, et al · 2024
Closest in time.
Deepseek-v3 technical report, 2024
DeepSeek-AI · 2024
Closest in time.
Codexgraph: Bridging large language models and code repositories via code graph databases, 2024
Xiangyan Liu, Bo Lan, Zhiyuan Hu, Yang Liu, Zhicheng Zhang, Fei Wang, Michael Shieh, and Wenmeng Zhou · 2024
Closest in time.
Yisen Xu, Feng Lin, Jinqiu Yang, Nikolaos Tsantalis, et al · 2025
Closest in time.
Mutation testing via iterative large language model-driven scientific debugging
Philipp Straubinger, Marvin Kreis, Stephan Lukasczyk, and Gordon Fraser · 2025
Closest in time.
14 most in-demand programming languages for 2025
iTransition · 2025
Closest in time.
When text embedding meets large language model: A comprehensive survey, 2025
Zhijie Nie, Zhangchi Feng, Mingxin Li, Cunwang Zhang, Yanzhao Zhang, Dingkun Long, and Richong Zhang · 2025
Closest in time.