Fetching the paper…
Reading the bibliography…
The expansion of the open source community and the rise of large language models have raised ethical and security concerns on the distribution of source code, such as misconduct on copyrighted code, distributions without proper licenses, or misuse of the code for malicious purposes.
A practical method for watermarking java programs
A. Monden, H. Iida, K. Matsumoto, K. Inoue, and K. Torii · 2000
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Software watermarking in the frequency domain: implementation, analysis, and attacks
Christian Collberg and Tapas Ranjan Sahoo · 2005
Earlier work this paper cites.
Words are not enough: sentence level natural language watermarking
Mercan Topkara, Umut Topkara, and Mikhail J Atallah · 2006
Earlier work this paper cites.
The hiding virtues of ambiguity: quantifiably resilient watermarking of natural language text through synonym substitutions
Umut Topkara, Mercan Topkara, and Mikhail J Atallah · 2006
Earlier work this paper cites.
Syntactic tools for text watermarking
Hasan M Meral, Emre Sevinc, Bülent Sankur, A Sumru Özsoy, and Tunga Güngör · 2007
Earlier work this paper cites.
Watermarking the outputs of structured prediction with an application in statistical machine translation
Ashish Venugopal, Jakob Uszkoreit, David Talbot, Franz Josef Och, and Juri Ganitkevitch · 2011
Earlier work this paper cites.
Improving neural networks by preventing co-adaptation of feature detectors
Geoffrey E Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, and Ruslan R Salakhutdinov · 2012
Earlier work this paper cites.
A survey of digital watermarking techniques, applications and attacks
Prabhishek Singh and Ramneet Singh Chadha · 2013
Earlier work this paper cites.
Function level control flow obfuscation for software security
Vivek Balachandran, Ng Wee Keong, and Sabu Emmanuel · 2014
Earlier work this paper cites.
Practical linguistic steganography using contextual synonym substitution and a novel vertex coding method
Ching-Yun Chang and Stephen Clark · 2014
Earlier work this paper cites.
Learning phrase representations using rnn encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Software plagiarism detection with birthmarks based on dynamic key instruction sequences
Zhenzhou Tian, Qinghua Zheng, Ting Liu, Ming Fan, Eryue Zhuang, and Zijiang Yang · 2015
Earlier work this paper cites.
On the naturalness of software
Abram Hindle, Earl T Barr, Mark Gabel, Zhendong Su, and Premkumar Devanbu · 2016
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2016
Earlier work this paper cites.
Hiding images in plain sight: Deep steganography
Shumeet Baluja · 2017
Earlier work this paper cites.
Hidden path: dynamic software watermarking based on control flow obfuscation
Zhe Chen, Chunfu Jia, and Donghui Xu · 2017
Earlier work this paper cites.
Generating steganographic images via adversarial training
Jamie Hayes and George Danezis · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Large-scale and language-oblivious code authorship identification
Mohammed Abuhamad, Tamer AbuHmed, Aziz Mohaisen, and DaeHun Nyang · 2018
Cited alongside, same era.
Software watermarking for java program based on method name encoding
Jianping Chen, Kui Li, Wanzhi Wen, Weixu Chen, and Chenxue Yan · 2018
Cited alongside, same era.
A review of text watermarking: theory, methods, and applications
Nurul Shamimi Kamaruddin, Amirrudin Kamsin, Lip Yee Por, and Hameedur Rahman · 2018
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2018
Cited alongside, same era.
Exception handling-based dynamic software watermarking
Yilong Wang, Daofu Gong, Bin Lu, Fei Xiang, and Fenlin Liu · 2018
Cited alongside, same era.
Learning-based recursive aggregation of abstract syntax trees for code clone detection
Lutz Büch and Artur Andrzejak · 2019
Softmark: Software watermarking via a binary function relocation
Honggoo Kang, Yonghwi Kwon, Sangjin Lee, and Hyungjoon Koo · 2021
Later among the works it cites.
Codexglue: A machine learning benchmark dataset for code understanding and generation
Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin Clement, Dawn Drain, Daxin Jiang, Duyu Tang, et al · 2021
Later among the works it cites.
Codet5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation
Yue Wang, Weishi Wang, Shafiq Joty, and Steven CH Hoi · 2021
Later among the works it cites.
Multi-lingual evaluation of code generation models
Ben Athiwaratkun, Sanjay Krishna Gouda, Zijian Wang, Xiaopeng Li, Yuchen Tian, Ming Tan, Wasi Uddin Ahmad, Shiqi Wang, Qing Sun, Mingyue Shang, et al · 2022
Later among the works it cites.
Natgen: generative pre-training by “naturalizing” source code
Saikat Chakraborty, Toufique Ahmed, Yangruibo Ding, Premkumar T Devanbu, and Baishakhi Ray · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Software watermarking: Progress and challenges
Ayan Dey, Sukriti Bhattacharya, and Nabendu Chaki · 2019
Cited alongside, same era.
Rethinking deep neural network ownership verification: Embedding passports to defeat ambiguity attacks
Lixin Fan, Kam Woh Ng, and Chee Seng Chan · 2019
Cited alongside, same era.
Codesearchnet challenge: Evaluating the state of semantic code search
Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt · 2019
Cited alongside, same era.
Code authorship attribution: Methods and challenges
Vaibhavi Kalgutkar, Ratinder Kaur, Hugo Gonzalez, Natalia Stakhanova, and Alina Matyukhina · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
Ropgen: Towards robust code authorship attribution via automatic coding style transformation
Zhen Li, Guenevere Chen, Chen Chen, Yayi Zou, and Shouhuai Xu · 2022
Later among the works it cites.
Coprotector: Protect open-source code against unauthorized training usage with data poisoning
Zhensu Sun, Xiaoning Du, Fu Song, Mingze Ni, and Li Li · 2022
Later among the works it cites.
You see what i want you to see: poisoning vulnerabilities in neural code search
Yao Wan, Shijie Zhang, Hongyu Zhang, Yulei Sui, Guandong Xu, Dezhong Yao, Hai Jin, and Lichao Sun · 2022
Later among the works it cites.
Bridging pre-trained models and downstream tasks for source code understanding
Deze Wang, Zhouyang Jia, Shanshan Li, Yue Yu, Yun Xiong, Wei Dong, and Xiangke Liao · 2022
Later among the works it cites.
Tracing text provenance via context-aware lexical substitution
Xi Yang, Jie Zhang, Kejiang Chen, Weiming Zhang, Zehua Ma, Feng Wang, and Nenghai Yu · 2022
Later among the works it cites.
Natural attack for pre-trained models of code
Zhou Yang, Jieke Shi, Junda He, and David Lo · 2022
Later among the works it cites.
100 million developers and counting, 2023
Thomas Dohmke · 2023
Closest in time.
Codeattack: Code-based adversarial attacks for pre-trained programming language models
Akshita Jha and Chandan K. Reddy · 2023
Closest in time.
How secure is code generated by chatgpt?
Raphaël Khoury, Anderson R Avila, Jacob Brunelle, and Baba Mamadou Camara · 2023
Closest in time.
A watermark for large language models
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein · 2023
Closest in time.
Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang · 2023
Closest in time.
Introducing chatgpt, 2022
OpenAI · 2023
Closest in time.
Temporary policy: Chatgpt is banned, 2023
OpenAI · 2023
Closest in time.
Robust multi-bit natural language watermarking through invariant features
KiYoon Yoo, Wonhyuk Ahn, Jiho Jang, and Nojun Kwak · 2023
Closest in time.