Fetching the paper…
Reading the bibliography…
Recent years have witnessed significant progress in developing deep learning-based models for automated code completion.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, et al · 1901
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
A survey of digital watermarking techniques, applications and attacks
Prabhishek Singh and Ramneet Singh Chadha. 2013 · 2013
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Code completion with statistical language models. In ACM SIGPLAN Conference on Programming Language Design and Implementation, PLDI ’14, Edinburgh, United Kingdom - June 09 - 11, 2014 . ACM, 419–428
Veselin Raychev, Martin T. Vechev, and Eran Yahav. 2014 · 2014
Earlier work this paper cites.
Gated graph sequence neural networks
Yujia Li, Daniel Tarlow, Marc Brockschmidt, and Richard Zemel. 2015 · 2015
Earlier work this paper cites.
Neural code completion
Chang Liu, Xin Wang, Richard Shin, Joseph E Gonzalez, and Dawn Song. 2016 · 2016
Earlier work this paper cites.
Membership Inference Attacks Against Machine Learning Models. In 2017 IEEE Symposium on Security and Privacy, SP 2017, San Jose, CA, USA, May 22-26, 2017 . IEEE Computer Society, 3–18
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. 2017 · 2017
Earlier work this paper cites.
Attention is all you need. In Proceedings of Advances in neural information processing systems . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Code Completion with Neural Attention and Pointer Networks. In IJCAI . 4159–4165
Jian Li, Yue Wang, Michael R. Lyu, and Irwin King. 2018 · 2018
Earlier work this paper cites.
Digital watermarking for deep neural networks
Yuki Nagai, Yusuke Uchida, Shigeyuki Sakazawa, and Shin’ichi Satoh. 2018 · 2018
Earlier work this paper cites.
Privacy risk in machine learning: Analyzing the connection to overfitting. In 2018 IEEE 31st computer security foundations symposium (CSF) . IEEE, 268–282
Samuel Yeom, Irene Giacomelli, Matt Fredrikson, and Somesh Jha. 2018 · 2018
Earlier work this paper cites.
CodeSearchNet Challenge: Evaluating the State of Semantic Code Search
Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt. 2019 · 2019
Earlier work this paper cites.
How Bad Can It Git? Characterizing Secret Leakage in Public GitHub Repositories. In 26th Annual Network and Distributed System Security Symposium, NDSS 2019, San Diego, California, USA, February 24-27, 2019
Michael Meli, Matthew R. McNiece, and Bradley Reaves. 2019 · 2019
Earlier work this paper cites.
Protection of bio medical iris image using watermarking and cryptography with WPT
R Mothi and M Karthikeyan. 2019 · 2019
Earlier work this paper cites.
ML-Leaks: Model and Data Independent Membership Inference Attacks and Defenses on Machine Learning Models. In 26th Annual Network and Distributed System Security Symposium, NDSS 2019, San Diego, California, USA, February 24-27, 2019
Ahmed Salem, Yang Zhang, Mathias Humbert, Pascal Berrang, Mario Fritz, and Michael Backes. 2019a · 2019
Earlier work this paper cites.
ML-Leaks: Model and Data Independent Membership Inference Attacks and Defenses on Machine Learning Models. In 26th Annual Network and Distributed System Security Symposium, NDSS 2019, San Diego, California, USA, February 24-27, 2019 . The Internet Society
Ahmed Salem, Yang Zhang, Mathias Humbert, Pascal Berrang, Mario Fritz, and Michael Backes. 2019b · 2019
Earlier work this paper cites.
Pythia: Ai-assisted code completion system. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . 2727–2735
Alexey Svyatkovskiy, Ying Zhao, Shengyu Fu, and Neel Sundaresan. 2019 · 2019
Cited alongside, same era.
Structural Language Models of Code. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event (Proceedings of Machine Learning Research, Vol. 119) . PMLR, 245–256
Uri Alon, Roy Sadaka, Omer Levy, and Eran Yahav. 2020 · 2020
Cited alongside, same era.
Sequence model design for code completion in the modern IDE
Gareth Ari Aye and Gail E Kaiser. 2020 · 2020
Cited alongside, same era.
Joint watermarking-encryption-JPEG-LS for medical image reliability control in encrypted and compressed domains
Sahar Haddad, Gouenou Coatrieux, Alexandre Moreau-Gaudry, and Michel Cozic. 2020 · 2020
Cited alongside, same era.
ReACC: A Retrieval-Augmented Code Completion Framework. In ACL . 6227–6240
Shuai Lu, Nan Duan, Hojae Han, Daya Guo, Seung-won Hwang, and Alexey Svyatkovskiy. 2022 · 2022
Later among the works it cites.
Coprotector: Protect open-source code against unauthorized training usage with data poisoning. In Proceedings of the ACM Web Conference 2022 . 652–660
Zhensu Sun, Xiaoning Du, Fu Song, Mingze Ni, and Li Li. 2022 · 2022
Later among the works it cites.
NaturalCC: An Open-Source Toolkit for Code Intelligence. In Proceedings of 44th International Conference on Software Engineering, Companion Volume . ACM
Yao Wan, Yang He, Zhangqian Bi, Jianguo Zhang, Yulei Sui, Hongyu Zhang, Kazuma Hashimoto, Hai Jin, Guandong Xu, Caiming Xiong, and Philip S. Yu. 2022 · 2022
Later among the works it cites.
A Systematic Evaluation of Large Language Models of Code
Frank F. Xu, Uri Alon, Graham Neubig, and Vincent J. Hellendoorn. 2022 · 2022
Later among the works it cites.
About Github Copilot telemetry
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Intellicode compose: Code generation using transformer. In ESEC/FSE . 1433–1443
Alexey Svyatkovskiy, Shao Kun Deng, Shengyu Fu, and Neel Sundaresan. 2020 · 2020
Cited alongside, same era.
Extracting Training Data from Large Language Models.. In USENIX Security Symposium , Vol. 6
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom B Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Cited alongside, same era.
Code Prediction by Feeding Trees to Transformers. In 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE) . 150–162
Seohyun Kim, Jinman Zhao, Yuchi Tian, and Satish Chandra. 2021 · 2021
Cited alongside, same era.
A survey of deep neural network watermarking techniques
Yue Li, Hongxia Wang, and Mauro Barni. 2021 · 2021
Cited alongside, same era.
Codexglue: A machine learning benchmark dataset for code understanding and generation
Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin Clement, Dawn Drain, Daxin Jiang, Duyu Tang, et al · 2021
Cited alongside, same era.
Systematic Evaluation of Privacy Risks of Machine Learning Models. In 30th USENIX Security Symposium, USENIX Security 2021, August 11-13, 2021 , Michael D. Bailey and Rachel Greenstadt (Eds.). USENIX Association, 2615–2632
Liwei Song and Prateek Mittal. 2021 · 2021
Cited alongside, same era.
Faketagger: Robust safeguards against deepfake dissemination via provenance tracking. In Proceedings of the 29th ACM International Conference on Multimedia . 3546–3555
Run Wang, Felix Juefei-Xu, Meng Luo, Yang Liu, and Lina Wang. 2021 · 2021
Cited alongside, same era.
Code Completion by Modeling Flattened Abstract Syntax Trees as Graphs. In Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Thirty-Third Conference on Innovative Applications of Artificial Intelligence, IAAI 2021, The Eleventh Symposium on Educational Advances in Artificial Intelligence, EAAI 2021, Virtual Event, February 2-9, 2021 . AAAI Press, 14015–14023
Yanlin Wang and Hui Li. 2021 · 2021
Cited alongside, same era.
2022 · 2023
Later among the works it cites.
StarCoder: may the source be with you!
Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, Qian Liu, Evgenii Zheltonozhskii, Terry Yue Zhuo, Thomas Wang, Olivier Dehaene, Mishig Davaadorj, Joel Lamy-Poirier, João Monteiro, Oleh Shliazhko, Nicolas Gontier, Nicholas Meade, Armel Zebaze, Ming-Ho Yee, Logesh Kumar Umapathi, Jian Zhu, Benjamin Lipkin, Muhtasham Oblokulov, Zhiruo Wang, Rudra Murthy, Jason Stillerman, Siva Sankalp Patel, Dmitry Abulkhanov, Marco Zocca, Manan Dey, Zhihan Zhang, Nour Fahmy, Urvashi Bhattacharyya, Wenhao Yu, Swayam Singh, Sasha Luccioni, Paulo Villegas, Maxim Kunakov, Fedor Zhdanov, Manuel Romero, Tony Lee, Nadav Timor, Jennifer Ding, Claire Schlesinger, Hailey Schoelkopf, Jan Ebert, Tri Dao, Mayank Mishra, Alex Gu, Jennifer Robinson, Carolyn Jane Anderson, Brendan Dolan-Gavitt, Danish Contractor, Siva Reddy, Daniel Fried, Dzmitry Bahdanau, Yacine Jernite, Carlos Muñoz Ferrandis, Sean Hughes, Thomas Wolf, Arjun Guha, Leandro von Werra, and Harm de Vries. 2023 · 2023
Later among the works it cites.
CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis
Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. 2023 · 2023
Later among the works it cites.
{ \{ CodexLeaks } \} : Privacy Leaks from Code Generation Language Models in { \{ GitHub } \} Copilot. In 32nd USENIX Security Symposium (USENIX Security 23) . 2133–2150
Liang Niu, Shujaat Mirza, Zayd Maradni, and Christina Pöpper. 2023 · 2023
Later among the works it cites.
A Novel Model Watermarking for Protecting Generative Adversarial Network
Tong Qiao, Yuyan Ma, Ning Zheng, Hanzhou Wu, Yanli Chen, Ming Xu, and Xiangyang Luo. 2023 · 2023
Later among the works it cites.
Code llama: Open foundation models for code
Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, et al · 2023
Later among the works it cites.
Code Llama: Open Foundation Models for Code
Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, Artyom Kozhevnikov, Ivan Evtimov, Joanna Bitton, Manish Bhatt, Cristian Canton Ferrer, Aaron Grattafiori, Wenhan Xiong, Alexandre Défossez, Jade Copet, Faisal Azhar, Hugo Touvron, Louis Martin, Nicolas Usunier, Thomas Scialom, and Gabriel Synnaeve. 2023 · 2023
Later among the works it cites.
CodeMark: Imperceptible Watermarking for Code Datasets against Neural Code Completion Models
Zhensu Sun, Xiaoning Du, Fu Song, and Li Li. 2023 · 2023
Later among the works it cites.
CodeT5+: Open Code Large Language Models for Code Understanding and Generation
Yue Wang, Hung Le, Akhilesh Deepak Gotmare, Nghi D. Q. Bui, Junnan Li, and Steven C. H. Hoi. 2023 · 2023
Later among the works it cites.
Gotcha! this model uses my code! evaluating membership leakage risks in code models
Zhou Yang, Zhipeng Zhao, Chenyu Wang, Jieke Shi, Dongsum Kim, Donggyun Han, and David Lo. 2023 · 2023
Later among the works it cites.
Counterfactual Memorization in Neural Language Models
Chiyuan Zhang, Daphne Ippolito, Katherine Lee, Matthew Jagielski, Florian Tramèr, and Nicholas Carlini. 2023 · 2023
Later among the works it cites.
Deep Learning for Code Generation: A Survey
Zhang Huangzhao, Zhang Kechi, Li Zhuo, Li Jia, Li Yongmin, Zhao Yunfei, Zhu Yuqi, Liu Fang, Li Ge, and Jin Zhi. 2024 · 2024
Closest in time.
Unveiling Memorization in Code Models. In 2024 IEEE/ACM 46th International Conference on Software Engineering (ICSE) . IEEE Computer Society, 856–856
Zhou Yang, Zhipeng Zhao, Chenyu Wang, Jieke Shi, Dongsun Kim, Donggyun Han, and David Lo. 2024 · 2024
Closest in time.