Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have demonstrated remarkable success in various natural language processing and software engineering tasks, such as code generation.
“Complementarity, f-score, and nlp evaluation,”
Leon Derczynski, · 2016
Earlier work this paper cites.
“Cyclomatic complexity,”
Christof Ebert, James Cain, Giuliano Antoniol, Steve Counsell, and Phillip Laplante, · 2016
Earlier work this paper cites.
“Comparing computational thinking development assessment scores with software complexity metrics,”
Jesús Moreno-León, Gregorio Robles, and Marcos Román-González, · 2016
Earlier work this paper cites.
“Cluster analysis to estimate the difficulty of programming problems,”
Chowdhury Md Intisar and Yutaka Watanobe, · 2018
Earlier work this paper cites.
“A systematic review on code clone detection,”
Qurat Ul Ain, Wasi Haider Butt, Muhammad Waseem Anwar, Farooque Azam, and Bilal Maqbool, · 2019
Earlier work this paper cites.
“Roberta: A robustly optimized bert pretraining approach,”
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov, · 2019
Earlier work this paper cites.
“Language models are few-shot learners,”
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al., · 2020
Earlier work this paper cites.
“CodeBERT: A pre-trained model for programming and natural languages,”
Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, and Ming Zhou, · 2020
Earlier work this paper cites.
“Can neural clone detection generalize to unseen functionalitiesf,”
Chenyao Liu, Zeqi Lin, Jian-Guang Lou, Lijie Wen, and Dongmei Zhang, · 2021
Earlier work this paper cites.
“Recent advances in natural language processing via large pre-trained language models: A survey,”
Bonan Min, Hayley Ross, Elior Sulem, Amir Pouran Ben Veyseh, Thien Huu Nguyen, Oscar Sainz, Eneko Agirre, Ilana Heinz, and Dan Roth, · 2021
Cited alongside, same era.
“Project codenet: A large-scale ai for code dataset for learning a diversity of coding tasks,”
Ruchir Puri, David S Kung, Geert Janssen, Wei Zhang, Giacomo Domeniconi, Vladmir Zolotov, Julian Dolby, Jie Chen, Mihir Choudhury, Lindsey Decker, et al., · 2021
Cited alongside, same era.
“Syncobert: Syntax-guided multi-modal contrastive pre-training for code representation,”
Xin Wang, Yasheng Wang, Fei Mi, Pingyi Zhou, Yao Wan, Xiao Liu, Li Li, Hao Wu, Jin Liu, and Xin Jiang, · 2021
Cited alongside, same era.
“Codexglue: A machine learning benchmark dataset for code understanding and generation,”
Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin Clement, Dawn Drain, Daxin Jiang, Duyu Tang, et al., · 2021
“A systematic evaluation of large language models of code,”
Frank F Xu, Uri Alon, Graham Neubig, and Vincent Josua Hellendoorn, · 2022
Later among the works it cites.
“Codexdb: Synthesizing code for query processing from natural language instructions using gpt-3 codex,”
Immanuel Trummer, · 2022
Later among the works it cites.
“Code generation tools (almost) for free? a study of few-shot, pre-trained language models on code,”
Patrick Bareiß, Beatriz Souza, Marcelo d’Amorim, and Michael Pradel, · 2022
Later among the works it cites.
“Retrieval-based prompt selection for code-related few-shot learning,”
Noor Nashid, Mifta Sintaha, and Ali Mesbah, · 2023
Later among the works it cites.
“Unveiling the potential of large language models in generating semantic and cross-language clones,”
Palash R Roy, Ajmain I Alam, Farouq Al-omari, Banani Roy, Chanchal K Roy, and Kevin A Schneider, · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Graphcode{bert}: Pre-training code representations with data flow,”
Daya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng, Duyu Tang, Shujie LIU, Long Zhou, Nan Duan, Alexey Svyatkovskiy, Shengyu Fu, Michele Tufano, Shao Kun Deng, Colin Clement, Dawn Drain, Neel Sundaresan, Jian Yin, Daxin Jiang, and Ming Zhou, · 2021
Cited alongside, same era.
“Deep learning application on code clone detection: A review of current knowledge,”
Maggie Lei, Hao Li, Ji Li, Namrata Aundhkar, and Dae-Kyoo Kim, · 2022
Cited alongside, same era.
“Generalizability of code clone detection on codebert,”
Tim Sonnekalb, Bernd Gruner, Clemens-Alexander Brust, and Patrick Mäder, · 2022
Cited alongside, same era.
“C4: Contrastive cross-language code clone detection,”
Chenning Tao, Qi Zhan, Xing Hu, and Xin Xia, · 2022
Cited alongside, same era.
“Identification and visualization of key topics in scientific publications with transformer-based language models and document clustering methods,”
Min-Hsien Weng, Shaoqun Wu, and Mark Dyer, · 2022
Cited alongside, same era.
Later among the works it cites.
“Towards understanding the capability of large language models on code clone detection: a survey,”
Shihan Dou, Junjie Shan, Haoxiang Jia, Wenhao Deng, Zhiheng Xi, Wei He, Yueming Wu, Tao Gui, Yang Liu, and Xuanjing Huang, · 2023
Later among the works it cites.
“Graph-based code semantics learning for efficient semantic code clone detection,”
Dongjin Yu, Quanxin Yang, Xin Chen, Jie Chen, and Yihang Xu, · 2023
Later among the works it cites.
“Improving cross-language code clone detection via code representation learning and graph neural networks,”
Nikita Mehrotra, Akash Sharma, Anmol Jindal, and Rahul Purandare, · 2023
Later among the works it cites.