Fetching the paper…
Reading the bibliography…
Source code authorship attribution is important in software forensics, plagiarism detection, and protecting software patch integrity.
Z. Li, Q. G. Chen, C. Chen, Y. Zou, and S. Xu, “RoPGen: Towards robust code authorship attribution via automatic coding style transformation,” in ICSE . ACM, 2022, pp. 1906–1918
1918
Earlier work this paper cites.
R. J. Leach, “Using metrics to evaluate student programs,” ACM SIGCSE Bull. , vol. 27, no. 2, pp. 41–43, 1995
1995
Earlier work this paper cites.
I. Krsul and E. H. Spafford, “Authorship analysis: identifying the author of a program,” Comput. Secur. , vol. 16, no. 3, pp. 233–257, 1997
1997
Earlier work this paper cites.
L. Prechelt, G. Malpohl, and M. Philippsen, “Finding plagiarisms among a set of programs with JPlag,” J. Univers. Comput. Sci. , vol. 8, no. 11, p. 1016, 2002
2002
Earlier work this paper cites.
S. Burrows and S. M. Tahaghoghi, “Source code authorship attribution using n-grams,” in ADCS , 2007, pp. 32–39
2007
Earlier work this paper cites.
M. Shevertalov, J. Kothari, E. Stehle, and S. Mancoridis, “On the use of discretized source code metrics for author identification,” in SSBSE . IEEE, 2009, pp. 69–78
2009
Earlier work this paper cites.
A. Caliskan-Islam, R. E. Harang, A. Liu, A. Narayanan, C. R. Voss, F. Yamaguchi, and R. Greenstadt, “De-anonymizing programmers via code stylometry,” in USENIX Security , J. Jung and T. Holz, Eds. USENIX Association, 2015, pp. 255–270
2015
Earlier work this paper cites.
S. Alrabaee, P. Shirani, M. Debbabi, and L. Wang, “On the feasibility of malware authorship attribution,” in FPS , ser. LNCS, F. Cuppens, L. Wang, N. Cuppens-Boulahia, N. Tawbi, and J. García-Alfaro, Eds., vol. 10128. Springer, 2016, pp. 256–272
2016
Earlier work this paper cites.
B. Alsulami, E. Dauber, R. E. Harang, S. Mancoridis, and R. Greenstadt, “Source code authorship attribution using long short-term memory based networks,” in ESORICS , ser. LNCS, S. N. Foley, D. Gollmann, and E. Snekkenes, Eds., vol. 10492. Springer, 2017, pp. 65–82
2017
Earlier work this paper cites.
M. Abuhamad, T. AbuHmed, A. Mohaisen, and D. Nyang, “Large-scale and language-oblivious code authorship identification,” in CCS , D. Lie, M. Mannan, M. Backes, and X. Wang, Eds. ACM, 2018, pp. 101–114
2018
Earlier work this paper cites.
V. Kalgutkar, R. Kaur, H. Gonzalez, N. Stakhanova, and A. Matyukhina, “Code authorship attribution: Methods and challenges,” ACM Comput. Surv. , vol. 52, no. 1, pp. 3:1–3:36, 2019
2019
Earlier work this paper cites.
E. Quiring, A. Maier, and K. Rieck, “Misleading authorship attribution of source code using adversarial learning,” in USENIX Security , N. Heninger and P. Traynor, Eds. USENIX Association, 2019, pp. 479–496
2019
Earlier work this paper cites.
M. Tufano, C. Watson, G. Bavota, M. D. Penta, M. White, and D. Poshyvanyk, “An empirical study on learning bug-fixing patches in the wild via neural machine translation,” ACM Trans. Softw. Eng. Methodol. , vol. 28, no. 4, pp. 19:1–19:29, 2019
2019
Earlier work this paper cites.
J. Vig, “A multiscale visualization of attention in the transformer model,” in ACL , M. R. Costa-jussà and E. Alfonseca, Eds. Association for Computational Linguistics, 2019, pp. 37–42
2019
Earlier work this paper cites.
M. Abuhamad, T. AbuHmed, D. Nyang, and D. Mohaisen, “Multi- χ \chi : Identifying multiple authors from source code files,” Proc. Priv. Enhancing Technol. , vol. 2020, no. 3, pp. 25–41, 2020
2020
Earlier work this paper cites.
Z. Feng, D. Guo, D. Tang, N. Duan, X. Feng, M. Gong, L. Shou, B. Qin, T. Liu, D. Jiang, and M. Zhou, “CodeBERT: A pre-trained model for programming and natural languages,” in EMNLP , ser. Findings of ACL, T. Cohn, Y. He, and Y. Liu, Eds., vol. EMNLP 2020. Association for Computational Linguistics, 2020, pp. 1536–1547
2020
Cited alongside, same era.
E. Bogomolov, V. Kovalenko, Y. Rebryk, A. Bacchelli, and T. Bryksin, “Authorship attribution of source code: a language-agnostic approach and applicability in software engineering,” in FSE , D. Spinellis, G. Gousios, M. Chechik, and M. D. Penta, Eds. ACM, 2021, pp. 932–944
2021
Cited alongside, same era.
2021
Cited alongside, same era.
D. Chicco, N. Tötsch, and G. Jurman, “The matthews correlation coefficient (MCC) is more reliable than balanced accuracy, bookmaker informedness, and markedness in two-class confusion matrix evaluation,” BioData Min. , vol. 14, no. 1, p. 13, 2021
Y. Deng, C. S. Xia, H. Peng, C. Yang, and L. Zhang, “Large language models are zero-shot fuzzers: Fuzzing deep-learning libraries via large language models,” in ISSTA , R. Just and G. Fraser, Eds. ACM, 2023, pp. 423–435
2023
Later among the works it cites.
2023
Later among the works it cites.
S. Choi and D. Mohaisen, “Attributing chatgpt-generated source codes,” IEEE Trans. Dependable Secur. Comput. , 2024
2024
Later among the works it cites.
Google, “Gemini,” https://ai.google.dev/gemini-api/docs , 2024, accessed: 2024-05-12
2024
Later among the works it cites.
OpenAI, “ChatGPT,” https://openai.com/blog/chatgpt , 2024, accessed: 2024-05-12
2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
A. Yuan, A. Coenen, E. Reif, and D. Ippolito, “Wordcraft: Story writing with large language models,” in IUI , G. Jacucci, S. Kaski, C. Conati, S. Stumpf, T. Ruotsalo, and K. Gajos, Eds. ACM, 2022, pp. 841–852
2022
Cited alongside, same era.
F. F. Xu, U. Alon, G. Neubig, and V. J. Hellendoorn, “A systematic evaluation of large language models of code,” in MAPS@PLDI , S. Chaudhuri and C. Sutton, Eds. ACM, 2022, pp. 1–10
2022
Cited alongside, same era.
N. Jain, S. Vaidyanath, A. S. Iyer, N. Natarajan, S. Parthasarathy, S. K. Rajamani, and R. Sharma, “Jigsaw: Large language models meet program synthesis,” in ICSE . ACM, 2022, pp. 1219–1231
2022
Cited alongside, same era.
S. Choi, R. Jang, D. Nyang, and D. Mohaisen, “Untargeted code authorship evasion with Seq2Seq transformation,” in CSoNet , ser. LNCS, M. H. Hà, X. Zhu, and M. T. Thai, Eds., vol. 14479. Springer, 2023, pp. 83–92
2023
Cited alongside, same era.
Wayne Xin Zhao et al., “A survey of large language models,” CoRR , vol. abs/2303.18223, 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
J. Leinonen, A. Hellas, S. Sarsa, B. N. Reeves, P. Denny, J. Prather, and B. A. Becker, “Using large language models to enhance programming error messages,” in SIGCSE , M. Doyle, B. Stephenson, B. Dorn, L. Soh, and L. Battestilli, Eds. ACM, 2023, pp. 563–569
2023
Cited alongside, same era.
Albert Q. Jiang et al., “Mistral 7b,” CoRR , vol. abs/2310.06825, 2023
2023
Cited alongside, same era.
Later among the works it cites.
Meta, “Llama,” https://llama.meta.com/ , 2024, accessed: 2024-05-12
2024
Later among the works it cites.
2024
Later among the works it cites.
OpenAI, “Document of chatgpt,” https://platform.openai.com/docs/model-index-for-researchers , 2024, accessed: 2024-05-12
2024
Later among the works it cites.
Y. Chang, X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y. Wang, W. Ye, Y. Zhang, Y. Chang, P. S. Yu, Q. Yang, and X. Xie, “A survey on evaluation of large language models,” ACM Trans. Intell. Syst. Technol. , vol. 15, no. 3, pp. 39:1–39:45, 2024
2024
Later among the works it cites.
Y. Liu, T. Le-Cong, R. Widyasari, C. Tantithamthavorn, L. Li, X. D. Le, and D. Lo, “Refining ChatGPT-generated code: Characterizing and mitigating code quality issues,” ACM Trans. Softw. Eng. Methodol. , vol. 33, no. 5, pp. 116:1–116:26, 2024
2024
Later among the works it cites.
C. Yan, M. H. Meng, F. Xie, and G. Bai, “Investigating documented privacy changes in android OS,” Proc. ACM Softw. Eng. , vol. 1, no. FSE, pp. 2701–2724, 2024
2024
Later among the works it cites.
Google, “Google code jam,” https://codingcompetitions.withgoogle.com/codejam/archive , 2024, accessed: 2024-05-12
2024
Later among the works it cites.
2024
Later among the works it cites.