Fetching the paper…
Reading the bibliography…
Understanding the functional (dis)-similarity of source code is significant for code modeling tasks such as software vulnerability and code clone detection.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Smote: Synthetic minority over-sampling technique
Nitesh V. Chawla, Kevin W. Bowyer, Lawrence O. Hall, and W. Philip Kegelmeyer. 2002 · 2002
Earlier work this paper cites.
MISIM: an end-to-end neural code similarity system
Fangke Ye, Shengtian Zhou, Anand Venkat, Ryan Marcus, Nesime Tatbul, Jesmin Jahan Tithi, Paul Petersen, Timothy G. Mattson, Tim Kraska, Pradeep Dubey, Vivek Sarkar, and Justin Gottschlich. 2020 · 2006
Earlier work this paper cites.
A hybrid approach (syntactic and textual) to clone detection
Marco Funaro, Daniele Braga, Alessandro Campi, and Carlo Ghezzi. 2010 · 2010
Earlier work this paper cites.
CLEAR: contrastive learning for sentence representation
Zhuofeng Wu, Sinong Wang, Jiatao Gu, Madian Khabsa, Fei Sun, and Hao Ma. 2020 · 2012
Earlier work this paper cites.
Towards a big data curated benchmark of inter-project code clones
Jeffrey Svajlenko, Judith F Islam, Iman Keivanloo, Chanchal K Roy, and Mohammad Mamun Mia. 2014 · 2014
Earlier work this paper cites.
Vulpecker: An automated vulnerability detection system based on code similarity analysis
Zhen Li, Deqing Zou, Shouhuai Xu, Hai Jin, Hanchao Qi, and Jie Hu. 2016 · 2016
Earlier work this paper cites.
Convolutional neural networks over tree structures for programming language processing
Lili Mou, Ge Li, Lu Zhang, Tao Wang, and Zhi Jin. 2016 · 2016
Earlier work this paper cites.
Vuddy: A scalable approach for vulnerable code clone discovery
Seulbae Kim, Seunghoon Woo, Heejo Lee, and Hakjoo Oh. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Learning to represent programs with graphs
Miltiadis Allamanis, Marc Brockschmidt, and Mahmoud Khademi. 2018 · 2018
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Earlier work this paper cites.
Vuldeepecker: A deep learning-based system for vulnerability detection
Zhen Li, Deqing Zou, Shouhuai Xu, Xinyu Ou, Hai Jin, Sujuan Wang, Zhijun Deng, and Yuyi Zhong. 2018b · 2018
Earlier work this paper cites.
A detection framework for semantic code clones and obfuscated code
Abdullah Sheneamer, Swarup Roy, and Jugal Kalita. 2018 · 2018
Earlier work this paper cites.
A systematic review on code clone detection
Qurat Ul Ain, Wasi Haider Butt, Muhammad Waseem Anwar, Farooque Azam, and Bilal Maqbool. 2019 · 2019
Earlier work this paper cites.
https://nvd.nist.gov/vuln/detail/CVE-2019-12730
CVE-2019-12730. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Devign: Effective vulnerability identification by learning comprehensive program semantics via graph neural networks
Yaqin Zhou, Shangqing Liu, Jingkai Siow, Xiaoning Du, and Yang Liu. 2019 · 2019
Cited alongside, same era.
Exploring software naturalness through neural language models
Luca Buratti, Saurabh Pujar, Mihaela Bornea, Scott McCarley, Yunhui Zheng, Gaetano Rossiello, Alessandro Morari, Jim Laredo, Veronika Thost, Yufan Zhuang, and Giacomo Domeniconi. 2020 · 2020
Cited alongside, same era.
Codit: Code editing with tree-based neural models
S. Chakraborty, Y. Ding, M. Allamanis, and B. Ray. 2020 · 2020
Cited alongside, same era.
Self-supervised contrastive learning for code retrieval and summarization via semantic-preserving transformations
Nghi D. Q. Bui, Yijun Yu, and Lingxiao Jiang. 2021 · 2021
Closest in time.
Deep learning based vulnerability detection: Are we there yet
Saikat Chakraborty, Rahul Krishna, Yangruibo Ding, and Baishakhi Ray. 2021 · 2021
Closest in time.
https://nvd.nist.gov/vuln/detail/CVE-2021-3449
CVE-2021-3449. 2021 · 2021
Closest in time.
https://nvd.nist.gov/vuln/detail/CVE-2021-38094
CVE-2021-38094. 2021 · 2021
Closest in time.
SimCSE: Simple contrastive learning of sentence embeddings
Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021 · 2021
Closest in time.
DeCLUTR: Deep contrastive learning for unsupervised textual representations
John Giorgi, Osvald Nitski, Bo Wang, and Gary Bader. 2021 · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020 · 2020
Cited alongside, same era.
https://nvd.nist.gov/vuln/detail/CVE-2020-24020
CVE-2020-24020. 2020 · 2020
Cited alongside, same era.
Patching as translation: the data and the metaphor
Yangruibo Ding, Baishakhi Ray, Devanbu Premkumar, and Vincent J. Hellendoorn. 2020 · 2020
Cited alongside, same era.
CodeBERT: A pre-trained model for programming and natural languages
Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, and Ming Zhou. 2020 · 2020
Cited alongside, same era.
Global relational models of source code
Vincent J. Hellendoorn, Charles Sutton, Rishabh Singh, Petros Maniatis, and David Bieber. 2020 · 2020
Cited alongside, same era.
Contrastive code representation learning
Paras Jain, Ajay Jain, Tianjun Zhang, Pieter Abbeel, Joseph E. Gonzalez, and Ion Stoica. 2020 · 2020
Cited alongside, same era.
Learning and evaluating contextual embedding of source code
Aditya Kanade, Petros Maniatis, Gogul Balakrishnan, and Kensen Shi. 2020 · 2020
Cited alongside, same era.
Graphcode{bert}: Pre-training code representations with data flow
Daya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng, Duyu Tang, Shujie LIU, Long Zhou, Nan Duan, Alexey Svyatkovskiy, Shengyu Fu, Michele Tufano, Shao Kun Deng, Colin Clement, Dawn Drain, Neel Sundaresan, Jian Yin, Daxin Jiang, and Ming Zhou. 2021 · 2021
Closest in time.
Treebert: A tree-based pre-trained model for programming language
Xue Jiang, Zhuoran Zheng, Chen Lyu, Liang Li, and Lei Lyu. 2021 · 2021
Closest in time.
Codexglue: A machine learning benchmark dataset for code understanding and generation
Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin B. Clement, Dawn Drain, Daxin Jiang, Duyu Tang, Ge Li, Lidong Zhou, Linjun Shou, Long Zhou, Michele Tufano, Ming Gong, Ming Zhou, Nan Duan, Neel Sundaresan, Shao Kun Deng, Shengyu Fu, and Shujie Liu. 2021 · 2021
Closest in time.
Cotext: Multi-task learning with code-text transformer
Long N. Phan, Hieu Tran, Daniel Le, Hieu Nguyen, James T. Anibal, Alec Peltekian, and Yanfang Ye. 2021 · 2021
Closest in time.
Contrastive learning with hard negative samples
Joshua David Robinson, Ching-Yao Chuang, Suvrit Sra, and Stefanie Jegelka. 2021 · 2021
Closest in time.
DOBF: A deobfuscation pre-training objective for programming languages
Baptiste Rozière, Marie-Anne Lachaux, Marc Szafraniec, and Guillaume Lample. 2021 · 2021
Closest in time.
Codet5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation
Yue Wang, Weishi Wang, Shafiq Joty, and Steven C.H. Hoi. 2021 · 2021
Closest in time.
Language-agnostic representation learning of source code from structure and context
Daniel Zügner, Tobias Kirschstein, Michele Catasta, Jure Leskovec, and Stephan Günnemann. 2021 · 2021
Closest in time.