Fetching the paper…
Reading the bibliography…
Code cloning, the duplication of code fragments, is common in software development.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
A learning algorithm for Boltzmann machines
David H Ackley, Geoffrey E Hinton, and Terrence J Sejnowski. 1985 · 1985
Earlier work this paper cites.
Substring matching for clone detection and change tracking.. In Proceedings of the 1994 International Conference on Software Maintenance (ICSM’94)
J Howard Johnson. 1994 · 1994
Earlier work this paper cites.
On Finding Duplication and Near-Duplication in Large Software Systems. In Proceedings of the Second Working Conference on Reverse Engineering (WCRE ’95) . IEEE Computer Society, USA, 86
B. S. Baker. 1995 · 1995
Earlier work this paper cites.
A language independent approach for detecting duplicated code. In Proceedings of the 1999 International Conference on Software Maintenance (ICSM’99)
Stéphane Ducasse, Matthias Rieger, and Serge Demeyer. 1999 · 1999
Earlier work this paper cites.
Using slicing to identify duplication in source code. In Proceedings of the 2001 International Static Analysis Symposium (ISAS’01)
Raghavan Komondoor and Susan Horwitz. 2001 · 2001
Earlier work this paper cites.
Identifying similar code with program dependence graphs. In Proceedings of the 8th Working Conference on Reverse Engineering (WCRE’01)
Jens Krinke. 2001 · 2001
Earlier work this paper cites.
CCFinder: a multilinguistic token-based code clone detection system for large scale source code
Toshihiro Kamiya, Shinji Kusumoto, and Katsuro Inoue. 2002 · 2002
Earlier work this paper cites.
An empirical study of code clone genealogies. In ESEC/FSE-13
Miryung Kim, Vibha Sazawal, David Notkin, and Gail C. Murphy. 2005 · 2005
Earlier work this paper cites.
Comparison and evaluation of clone detection tools
Stefan Bellon, Rainer Koschke, Giulio Antoniol, Jens Krinke, and Ettore Merlo. 2007 · 2007
Earlier work this paper cites.
Deckard: scalable and accurate tree-based detection of code clones. In Proceedings of the 29th International Conference on Software Engineering (ICSE’07)
Lingxiao Jiang, Ghassan Misherghi, Zhendong Su, and Stephane Glondu. 2007 · 2007
Earlier work this paper cites.
A survey on software clone detection research
Chanchal Kumar Roy and James R Cordy. 2007 · 2007
Earlier work this paper cites.
NICAD: accurate detection of near-miss intentional clones using flexible pretty-printing and code normalization. In Proceedings of the 2008 International Conference on Program Comprehension (ICPC’08)
Chanchal K Roy and James R Cordy. 2008 · 2008
Earlier work this paper cites.
Incremental clone detection. In Proceedings of the 2009 European Conference on Software Maintenance and Reengineering (ECSMR’09)
Nils Göde and Rainer Koschke. 2009 · 2009
Earlier work this paper cites.
Index-based code clone detection: incremental, distributed, scalable. In 2010 IEEE International Conference on Software Maintenance . IEEE, 1–9
Benjamin Hummel, Elmar Juergens, Lars Heinemann, and Michael Conradt. 2010 · 2010
Earlier work this paper cites.
Haskell Clone Detection using Pattern Comparing Algorithm. In Proceedings of the 13th International Conference on Engineering of Modern Electric Systems (EMES)
Sergej Chodarev, Emilia Pietrikova, and Jan Kollar. 2015 · 2015
Earlier work this paper cites.
Code Clones Detection Using Machine Learning Technique: Support Vector Machine. In Proceedings of the 2016 IEEE INTERNATIONAL CONFERENCE ON COMPUTING, COMMUNICATION AND AUTOMATION (ICCCA) . 299–303
Shruti Jadon. 2016 · 2016
Earlier work this paper cites.
SourcererCC: scaling code clone detection to big code. In Proceedings of the 38th International Conference on Software Engineering (ICSE’16)
Hitesh Sajnani, Vaibhav Saini, Jeffrey Svajlenko, Chanchal K Roy, and Cristina V Lopes. 2016 · 2016
Earlier work this paper cites.
Controlling Linguistic Style Aspects in Neural Language Generation. In Proceedings of the Workshop on Stylistic Variation . Association for Computational Linguistics, Copenhagen, Denmark, 94–104
Jessica Ficler and Yoav Goldberg. 2017 · 2017
Earlier work this paper cites.
Vuddy: A scalable approach for vulnerable code clone discovery. In 2017 IEEE Symposium on Security and Privacy (SP) . IEEE, 595–614
Seulbae Kim, Seunghoon Woo, Heejo Lee, and Hakjoo Oh. 2017 · 2017
Earlier work this paper cites.
Cclearner: a deep learning-based clone detection approach. In Proceedings of the 2017 International Conference on Software Maintenance and Evolution (ICSME’17)
Liuqing Li, He Feng, Wenjie Zhuang, Na Meng, and Barbara Ryder. 2017 · 2017
Earlier work this paper cites.
A Comparison Among ARIMA, BP-NN, and MOGA-NN for Software Clone Evolution Prediction
Jayadeep Pati, Babloo Kumar, Devesh Manjhi, and K. K. Shukla. 2017 · 2017
Earlier work this paper cites.
Using Compilation/Decompilation to Enhance Clone Detection. In Proceedings of the 11th IEEE International Workshop on Software Clones (IWSC) . 8–14
Chaiyong Ragkhitwetsagul and Jens Krinke. 2017 · 2017
Earlier work this paper cites.
CCSharp: an efficient three-phase code clone detector using modified pdgs. In Proceedings of the 24th Asia-Pacific Software Engineering Conference (APSEC’17)
Min Wang, Pengcheng Wang, and Yun Xu. 2017 · 2017
Earlier work this paper cites.
Supervised deep features for software functional clone detection by exploiting lexical and syntactical information in source code. In Proceedings of the 2017 International Joint Conferences on Artificial Intelligence (IJCAI’17)
Huihui Wei and Ming Li. 2017 · 2017
Earlier work this paper cites.
Detecting Java Code Clones with Multi-Granularities Based on Bytecode. In Proceedings of the 41st IEEE Annual Computer Software and Applications Conference (COMPSAC) . 317–326
Dongjin Yu, Jie Wang, Qing Wu, Jiazha Yang, Jiaojiao Wang, Wei Yang, and Wei Yan. 2017 · 2017
Earlier work this paper cites.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin. 2018 · 2018
Earlier work this paper cites.
Software systems at risk: An empirical study of cloned vulnerabilities in practice
Seulbae Kim and Heejo Lee. 2018 · 2018
Cited alongside, same era.
Oreo: detection of clones in the twilight zone. In Proceedings of the 26th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (FSE’18)
Vaibhav Saini, Farima Farmahinifarahani, Yadong Lu, Pierre Baldi, and Cristina V Lopes. 2018 · 2018
Cited alongside, same era.
CCAligner: a token based large-gap clone detector. In Proceedings of the 40th International Conference on Software Engineering (ICSE’18)
Pengcheng Wang, Jeffrey Svajlenko, Yanzhao Wu, Yun Xu, and Chanchal K Roy. 2018 · 2018
Cited alongside, same era.
Deepsim: deep learning code functional similarity. In Proceedings of the 26th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (FSE’18)
Gang Zhao and Jeff Huang. 2018 · 2018
Cited alongside, same era.
Deep learning application on code clone detection: A review of current knowledge
Maggie Lei, Hao Li, Ji Li, Namrata Aundhkar, and Dae-Kyoo Kim. 2022 · 2022
Later among the works it cites.
On the advance of making language models better reasoners
Yifei Li, Zeqi Lin, Shizhuo Zhang, Qiang Fu, Bei Chen, Jian-Guang Lou, and Weizhu Chen. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of NAACL-HLT . 4171–4186
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
How to fine-tune bert for text classification?. In Chinese Computational Linguistics: 18th China National Conference, CCL 2019, Kunming, China, October 18–20, 2019, Proceedings 18 . Springer, 194–206
Chi Sun, Xipeng Qiu, Yige Xu, and Xuanjing Huang. 2019 · 2019
Cited alongside, same era.
A novel neural source code representation based on abstract syntax tree. In Proceedings of the 41st International Conference on Software Engineering (ICSE’19)
Jian Zhang, Xu Wang, Hongyu Zhang, Hailong Sun, Kaixuan Wang, and Xudong Liu. 2019 · 2019
Cited alongside, same era.
BigCloneBench
2020 · 2020
Cited alongside, same era.
Codebert: A pre-trained model for programming and natural languages
Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, et al · 2020
Cited alongside, same era.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
Detecting Semantic Code Clones by Building AST-based Markov Chains Model. In Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering (ASE’22)
Yueming Wu, Siyue Feng, Deqing Zou, and Hai Jin. 2022 · 2022
Later among the works it cites.
A survey of automatic source code summarization
Chunyan Zhang, Junchao Wang, Qinglei Zhou, Ting Xu, Ke Tang, Hairen Gui, and Fudong Liu. 2022 · 2022
Later among the works it cites.
Falcon-40B: an open large language model with state-of-the-art performance
Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Merouane Debbah, Etienne Goffinet, Daniel Heslow, Julien Launay, Quentin Malartic, et al · 2023
Closest in time.
Large Language Models are Zero-Shot Fuzzers: Fuzzing Deep-Learning Libraries via Large Language Models. In Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis . 423–435
Yinlin Deng, Chunqiu Steven Xia, Haoran Peng, Chenyuan Yang, and Lingming Zhang. 2023 · 2023
Closest in time.
Prompting Is All Your Need: Automated Android Bug Replay with Large Language Models
Sidong Feng and Chunyang Chen. 2023 · 2023
Closest in time.
CMMLU: Measuring massive multitask language understanding in Chinese
Haonan Li, Yixuan Zhang, Fajri Koto, Yifei Yang, Hai Zhao, Yeyun Gong, Nan Duan, and Timothy Baldwin. 2023b · 2023
Closest in time.
StarCoder: may the source be with you!
Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, et al · 2023
Closest in time.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023b · 2023
Closest in time.
Jailbreaking chatgpt via prompt engineering: An empirical study
Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, and Yang Liu. 2023a · 2023
Closest in time.
Recent Advances in Natural Language Processing via Large Pre-Trained Language Models: A Survey
Bonan Min, Hayley Ross, Elior Sulem, Amir Pouran Ben Veyseh, Thien Huu Nguyen, Oscar Sainz, Eneko Agirre, Ilana Heintz, and Dan Roth. 2023 · 2023
Closest in time.
GPT-4 Technical Report
OpenAI. 2023 · 2023
Closest in time.
Alpaca: A strong, replicable instruction-following model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023a · 2023
Closest in time.
Introducing MPT-30B: Raising the bar for open-source foundation models
MosaicML NLP Team. 2023 · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Closest in time.
Creating a Coding Assistant with StarCoder
Lewis Tunstall, Nathan Lambert, Nazneen Rajani, Edward Beeching, Teven Le Scao, Leandro von Werra, Sheon Han, Philipp Schmid, and Alexander Rush. 2023 · 2023
Closest in time.
Visionllm: Large language model is also an open-ended decoder for vision-centric tasks
Wenhai Wang, Zhe Chen, Xiaokang Chen, Jiannan Wu, Xizhou Zhu, Gang Zeng, Ping Luo, Tong Lu, Jie Zhou, Yu Qiao, et al · 2023
Closest in time.
Codet5+: Open code large language models for code understanding and generation
Yue Wang, Hung Le, Akhilesh Deepak Gotmare, Nghi DQ Bui, Junnan Li, and Steven CH Hoi. 2023b · 2023
Closest in time.
Multimodal chain-of-thought reasoning in language models
Zhuosheng Zhang, Aston Zhang, Mu Li, Hai Zhao, George Karypis, and Alex Smola. 2023 · 2023
Closest in time.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al · 2023
Closest in time.
Judging LLM-as-a-judge with MT-Bench and Chatbot Arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2023
Closest in time.
Secrets of RLHF in Large Language Models Part I: PPO
Rui Zheng, Shihan Dou, Songyang Gao, Wei Shen, Binghai Wang, Yan Liu, Senjie Jin, Qin Liu, Limao Xiong, Lu Chen, et al · 2023
Closest in time.