Fetching the paper…
Reading the bibliography…
Code datasets are of immense value for training neural-network-based code completion models, where companies or organizations have made substantial investments to establish and process these datasets.
The generalisation of student’s problems when several different population variances are involved
B. L. Welch. 1947 · 1947
Earlier work this paper cites.
The Mechanical Evaluation of Expressions
Peter J. Landin. 1964 · 1964
Earlier work this paper cites.
Method and system for generating and auditing a signature for a computer program
Robert I Davidson and Nathan Myhrvold. 1996 · 1996
Earlier work this paper cites.
A practical method for watermarking Java programs
Akito Monden, Hajimu Iida, Ken ichi Matsumoto, Koji Torii, and Katsuro Inoue. 2000 · 2000
Earlier work this paper cites.
A Method for Watermarking Java Programs via Opaque Predicates
Geneviève Arboit. 2002 · 2002
Earlier work this paper cites.
Bleu: a Method for Automatic Evaluation of Machine Translation. In ACL
Kishore Papineni, S. Roukos, T. Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Software watermarking via assembly code transformations
Smita Thaker. 2004 · 2004
Earlier work this paper cites.
An Evaluation of Static Java Bytecode Watermarking
Sebastian Danicic and James Alexander George Hamilton. 2010 · 2010
Earlier work this paper cites.
A survey of static software watermarking
James Alexander George Hamilton and Sebastian Danicic. 2011 · 2011
Earlier work this paper cites.
An Efficient Software Watermark by Equation Reordering and FDOS. In SocProS
B. K. Sharma, R. P. Agarwal, and Raghuraj Singh. 2011 · 2011
Earlier work this paper cites.
Targeted backdoor attacks on deep learning systems using data poisoning
Xinyun Chen, Chang Liu, Bo Li, Kimberly Lu, and Dawn Song. 2017 · 2017
Earlier work this paper cites.
BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
Tianyu Gu, Brendan Dolan-Gavitt, and S. Garg. 2017 · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Turning Your Weakness Into a Strength: Watermarking Deep Neural Networks by Backdooring. In USENIX Security Symposium
Yossi Adi, Carsten Baum, Moustapha Cissé, Benny Pinkas, and Joseph Keshet. 2018 · 2018
Earlier work this paper cites.
Public Git Archive: A Big Code Dataset for All
Vadim Markovtsev and Waren Long. 2018 · 2018
Earlier work this paper cites.
Poison Frogs! Targeted Clean-Label Poisoning Attacks on Neural Networks. In NeurIPS
A. Shafahi, W. R. Huang, Mahyar Najibi, O. Suciu, Christoph Studer, T. Dumitras, and T. Goldstein. 2018 · 2018
Earlier work this paper cites.
Spectral Signatures in Backdoor Attacks. In NeurIPS
Brandon Tran, Jerry Li, and A. Madry. 2018 · 2018
Earlier work this paper cites.
Detecting Backdoor Attacks on Deep Neural Networks by Activation Clustering
Bryant Chen, Wilka Carvalho, Nathalie Baracaldo, Heiko Ludwig, Ben Edwards, Taesung Lee, Ian Molloy, and B. Srivastava. 2019 · 2019
Earlier work this paper cites.
CodeSearchNet Challenge: Evaluating the State of Semantic Code Search
Hamel Husain, Hongqi Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt. 2019 · 2019
Cited alongside, same era.
Xmark: Dynamic Software Watermarking Using Collatz Conjecture
Haoyu Ma, Chunfu Jia, Shijia Li, Wantong Zheng, and Dinghao Wu. 2019 · 2019
Cited alongside, same era.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Digital Watermarking For Protecting Audio Classification Datasets
Wan Soo Kim and Kyogu Lee. 2020 · 2020
Cited alongside, same era.
Open-sourced Dataset Protection via Backdoor Watermarking
Yiming Li, Zi-Mou Zhang, Jiawang Bai, Baoyuan Wu, Yong Jiang, and Shutao Xia. 2020 · 2020
Cited alongside, same era.
Yue Wang, Weishi Wang, Shafiq R. Joty, and Steven C. H. Hoi. 2021 · 2021
Later among the works it cites.
A Targeted Attack on Black-Box Neural Machine Translation with Parallel Data Poisoning
Changming Xu, Jun Wang, Yuqing Tang, Francisco Guzmán, Benjamin I. P. Rubinstein, and Trevor Cohn. 2021 · 2021
Later among the works it cites.
Robust Black-box Watermarking for Deep Neural Network using Inverse Document Frequency
Mohammad Mehdi Yadollahi, Farzaneh Shoeleh, Sajjad Dadkhah, and Ali A. Ghorbani. 2021 · 2021
Later among the works it cites.
Challenging Machine Learning-based Clone Detectors via Semantic-preserving Code Transformations
Weiwei Zhang, Shengjian Guo, Hongyu Zhang, Yulei Sui, Yinxing Xue, and Yun Xu. 2021 · 2021
Later among the works it cites.
aiXcoder
2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Goutham Ramakrishnan and Aws Albarghouthi. 2020 · 2020
Cited alongside, same era.
You Autocomplete Me: Poisoning Vulnerabilities in Neural Code Completion
R. Schuster, Congzheng Song, Eran Tromer, and Vitaly Shmatikov. 2020 · 2020
Cited alongside, same era.
STRATA: Simple, Gradient-Free Attacks for Models of Code
Jacob M. Springer, Bryn Reinstadler, and Una-May O’Reilly. 2020 · 2020
Cited alongside, same era.
A Survey on Deep Learning for Software Engineering
Yanming Yang, Xin Xia, David Lo, and John C. Grundy. 2020 · 2020
Cited alongside, same era.
Adversarial examples for models of code
Noam Yefet, Uri Alon, and Eran Yahav. 2020 · 2020
Cited alongside, same era.
Generating Adversarial Examples for Holding Robustness of Source Code Processing Models. In AAAI
Huangzhao Zhang, Zhuo Li, Ge Li, L. Ma, Yang Liu, and Zhi Jin. 2020 · 2020
Cited alongside, same era.
Clean-Label Backdoor Attacks on Video Recognition Models
Shihao Zhao, Xingjun Ma, X. Zheng, J. Bailey, Jingjing Chen, and Yugang Jiang. 2020 · 2020
Cited alongside, same era.
Code faster with AI completions | TabNine
2022 · 2022
Later among the works it cites.
GitHub Copilot · Your AI pair programmer
2022a · 2022
Later among the works it cites.
Tree-sitter - Introduction
2022 · 2022
Later among the works it cites.
The Stack: 3 TB of permissively licensed source code
Denis Kocetkov, Raymond Li, Loubna Ben Allal, Jia Li, Chenghao Mou, Carlos Muñoz Ferrandis, Yacine Jernite, Margaret Mitchell, Sean Hughes, Thomas Wolf, Dzmitry Bahdanau, Leandro von Werra, and Harm de Vries. 2022 · 2022
Later among the works it cites.
On the Effectiveness of Dataset Watermarking in Adversarial Settings
Buse Gul Atli Tekgul and N. Asokan. 2022 · 2022
Later among the works it cites.
Natural Attack for Pre-trained Models of Code
Zhou Yang, Jieke Shi, Junda He, and David Lo. 2022 · 2022
Later among the works it cites.
How is the data in Copilot for Individuals used and shared?
2022b · 2023
Closest in time.
ML-powered coding companion – Amazon CodeWhisperer – Amazon Web Services
2022 · 2023
Closest in time.
Where did AWS obtain the training data to build this service?
2022 · 2023
Closest in time.
CodeMark
2023 · 2023
Closest in time.
Stack Overflow Will Charge AI Giants for Training Data
2023 · 2023
Closest in time.
StarCoder: may the source be with you!
Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, Qian Liu, Evgenii Zheltonozhskii, Terry Yue Zhuo, Thomas Wang, Olivier Dehaene, Mishig Davaadorj, Joel Lamy-Poirier, João Monteiro, Oleh Shliazhko, Nicolas Gontier, Nicholas Meade, Armel Zebaze, Ming-Ho Yee, Logesh Kumar Umapathi, Jian Zhu, Benjamin Lipkin, Muhtasham Oblokulov, Zhiruo Wang, Rudra Murthy, Jason Stillerman, Siva Sankalp Patel, Dmitry Abulkhanov, Marco Zocca, Manan Dey, Zhihan Zhang, Nourhan Fahmy, Urvashi Bhattacharyya, W. Yu, Swayam Singh, Sasha Luccioni, Paulo Villegas, Maxim Kunakov, Fedor Zhdanov, Manuel Romero, Tony Lee, Nadav Timor, Jennifer Ding, Claire Schlesinger, Hailey Schoelkopf, Jana Ebert, Tri Dao, Mayank Mishra, Alexander Gu, Jennifer Robinson, Carolyn Jane Anderson, Brendan Dolan-Gavitt, Danish Contractor, Siva Reddy, Daniel Fried, Dzmitry Bahdanau, Yacine Jernite, Carlos Muñoz Ferrandis, Sean M. Hughes, Thomas Wolf, Arjun Guha, Leandro von Werra, and Harm de Vries. 2023 · 2023
Closest in time.