Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have demonstrated strong capabilities in various code intelligence tasks.
Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics . 311–318
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries. In Text summarization branches out . 74–81
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization . 65–72
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Drebin: Effective and explainable detection of android malware in your pocket. In Proceedings of the Network and Distributed Systems Security Symposium (NDSS) , Vol. 14. 23–26
Daniel Arp, Michael Spreitzenbarth, Malte Hubner, Hugo Gascon, Konrad Rieck, and CERT Siemens. 2014 · 2014
Earlier work this paper cites.
Better Malware Ground Truth: Techniques for Weighting Anti-Virus Vendor Labels. In Proceedings of the 8th ACM Workshop on Artificial Intelligence and Security, AISec 2015, Denver, Colorado, USA, October 16, 2015
Alex Kantchelian, Michael Carl Tschantz, Sadia Afroz, Brad Miller, Vaishaal Shankar, Rekha Bachwani, Anthony D. Joseph, and J. Doug Tygar. 2015 · 2015
Earlier work this paper cites.
Binary code is not easy. In Proceedings of the 25th International Symposium on Software Testing and Analysis . 24–35
Xiaozhu Meng and Barton P Miller. 2016 · 2016
Earlier work this paper cites.
Lightgbm: A highly efficient gradient boosting decision tree
Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie-Yan Liu. 2017 · 2017
Earlier work this paper cites.
MamaDroid: Detecting Android Malware by Building Markov Chains of Behavioral Models. In Proceedings of the Network and Distributed Systems Security Symposium (NDSS)
E Mariconti, L Onwuzurike, P Andriotis, E De Cristofaro, G Ross, and G Stringhini. 2017 · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Codesearchnet challenge: Evaluating the state of semantic code search
Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt. 2019 · 2019
Earlier work this paper cites.
{ \{ TESSERACT } \} : Eliminating experimental bias in malware classification across space and time. In 28th USENIX security symposium (USENIX Security 19) . 729–746
Feargus Pendlebury, Fabio Pierazzi, Roberto Jordaney, Johannes Kinder, and Lorenzo Cavallaro. 2019 · 2019
Earlier work this paper cites.
Evaluating explanation without ground truth in interpretable machine learning
Fan Yang, Mengnan Du, and Xia Hu. 2019 · 2019
Earlier work this paper cites.
Androguard
Anthony Desnos. [n. d.] · 2020
Earlier work this paper cites.
Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation Metrics. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . 4984–4997
Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2020 · 2020
Earlier work this paper cites.
Unified Pre-training for Program Understanding and Generation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . 2655–2668
Wasi Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2021 · 2021
Earlier work this paper cites.
DeepReflect: Discovering Malicious Functionality through Binary Reconstruction. In 30th USENIX Security Symposium (USENIX Security 21)
Evan Downing, Yisroel Mirsky, Kyuhong Park, and Wenke Lee. 2021 · 2021
Earlier work this paper cites.
Automatic detection of five api documentation smells: Practitioners’ perspectives. In 2021 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER) . IEEE, 318–329
Junaed Younus Khan, Md Tawkat Islam Khondaker, Gias Uddin, and Anindya Iqbal. 2021 · 2021
Earlier work this paper cites.
CodeXGLUE : A Machine Learning Benchmark Dataset for Code Understanding and Generation. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1)
Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin Clement, Dawn Drain, Daxin Jiang, Duyu Tang, et al · 2021
Cited alongside, same era.
Towards Accurate Labeling of Android Apps for Reliable Malware Detection. In Proceedings of the Eleventh ACM Conference on Data and Application Security and Privacy (CODASPY ’21)
Aleieldin Salem. 2021 · 2021
Cited alongside, same era.
CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . 8696–8708
Yue Wang, Weishi Wang, Shafiq Joty, and Steven CH Hoi. 2021 · 2021
Cited alongside, same era.
{ \{ RE-Mind } \} : a first look inside the mind of a reverse engineer. In 31st USENIX Security Symposium (USENIX Security 22) . 2727–2745
Alessandro Mantovani, Simone Aonzo, Yanick Fratantonio, and Davide Balzarotti. 2022 · 2022
ReCode: Robustness Evaluation of Code Generation Models. In The 61st Annual Meeting Of The Association For Computational Linguistics
Shiqi Wang, Zheng Li, Haifeng Qian, Chenghao Yang, Zijian Wang, Mingyue Shang, Varun Kumar, Samson Tan, Baishakhi Ray, Parminder Bhatia, et al · 2023
Later among the works it cites.
Codegeex: A pre-trained model for code generation with multilingual benchmarking on humaneval-x. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 5673–5684
Qinkai Zheng, Xiao Xia, Xu Zou, Yuxiao Dong, Shan Wang, Yufei Xue, Lei Shen, Zihan Wang, Andi Wang, Yang Li, et al · 2023
Later among the works it cites.
Codeplan: Repository-level coding using llms and planning
Ramakrishna Bairi, Atharv Sonwane, Aditya Kanade, Arun Iyer, Suresh Parthasarathy, Sriram Rajamani, B Ashok, and Shashank Shet. 2024 · 2024
Later among the works it cites.
Teaching large language models to self-debug. In The Twelfth International Conference on Learning Representations
Xinyun Chen, Maxwell Lin, Nathanael Schärli, and Denny Zhou. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
MalRadar: Demystifying Android malware in the new era
Liu Wang, Haoyu Wang, Ren He, Ran Tao, Guozhu Meng, Xiapu Luo, and Xuanzhe Liu. 2022 · 2022
Cited alongside, same era.
Finetuned Language Models are Zero-Shot Learners. In International Conference on Learning Representations
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. 2022 · 2022
Cited alongside, same era.
Extending source code pre-trained language models to summarise decompiled binaries. In 2023 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER) . IEEE, 260–271
Ali Al-Kaswan, Toufique Ahmed, Maliheh Izadi, Anand Ashok Sawant, Premkumar Devanbu, and Arie van Deursen. 2023 · 2023
Cited alongside, same era.
Humans vs. Machines in Malware Classification. In 32nd USENIX Security Symposium (USENIX Security 23)
Simone Aonzo, Yufei Han, Alessandro Mantovani, and Davide Balzarotti. 2023 · 2023
Cited alongside, same era.
Finer: Enhancing state-of-the-art classifiers with feature attribution to facilitate security analysis. In Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security . 416–430
Yiling He, Jian Lou, Zhan Qin, and Kui Ren. 2023b · 2023
Cited alongside, same era.
Mohammad Abdullah Matin Khan, M Saiful Bari, Xuan Long Do, Weishi Wang, Md Rizwan Parvez, and Shafiq Joty. 2023 · 2023
Cited alongside, same era.
The Stack: 3 TB of permissively licensed source code
Denis Kocetkov, Raymond Li, LI Jia, Chenghao Mou, Yacine Jernite, Margaret Mitchell, Carlos Muñoz Ferrandis, Sean Hughes, Thomas Wolf, Dzmitry Bahdanau, et al · 2023
Cited alongside, same era.
Cctest: Testing and repairing code completion systems. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE) . IEEE, 1238–1250
Zongjie Li, Chaozheng Wang, Zhibo Liu, Haoxuan Wang, Dong Chen, Shuai Wang, and Cuiyun Gao. 2023b · 2023
Cited alongside, same era.
Large Language Models for Code Analysis: Do LLMs Really Do Their Job?. In 33rd USENIX Security Symposium (USENIX Security 24) . USENIX Association, Philadelphia, PA, 829–846
Chongzhou Fang, Ning Miao, Shaurya Srivastav, Jialin Liu, Ruoyu Zhang, Ruijie Fang, Asmita, Ryan Tsang, Najmeh Nazari, Han Wang, and Houman Homayoun. 2024 · 2024
Later among the works it cites.
DeepSeek-Coder: When the Large Language Model Meets Programming–The Rise of Code Intelligence
Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Yu Wu, YK Li, et al · 2024
Later among the works it cites.
DREAM: Combating Concept Drift with Explanatory Detection and Adaptation in Malware Classification
Yiling He, Junchi Lei, Zhan Qin, and Kui Ren. 2024 · 2024
Later among the works it cites.
Enhancing Android Malware Detection: The Influence of ChatGPT on Decision-centric Task
Yao Li, Sen Fang, Tao Zhang, and Haipeng Cai. 2024 · 2024
Later among the works it cites.
How Far Have We Gone in Binary Code Understanding Using Large Language Models. In 2024 IEEE International Conference on Software Maintenance and Evolution (ICSME) . IEEE, 1–12
Xiuwei Shang, Shaoyin Cheng, Guoqiang Chen, Yanming Zhang, Li Hu, Xiao Yu, Gangyang Li, Weiming Zhang, and Nenghai Yu. 2024 · 2024
Later among the works it cites.
DebugBench: Evaluating Debugging Capability of Large Language Models. In Findings of the Association for Computational Linguistics ACL 2024 . 4173–4198
Runchu Tian, Yining Ye, Yujia Qin, Xin Cong, Yankai Lin, Yinxu Pan, Yesai Wu, Hui Haotian, Liu Weichuan, Zhiyuan Liu, et al · 2024
Later among the works it cites.
Tapi: Towards target-specific and adversarial prompt injection against code llms
Yuchen Yang, Hongwei Yao, Bingrun Yang, Yiling He, Yiming Li, Tianwei Zhang, Zhan Qin, and Kui Ren. 2024 · 2024
Later among the works it cites.
LAMD: Context-driven Android Malware Detection and Classification with LLMs
Xingzhi Qian, Xinran Zheng, Yiling He, Shuo Yang, and Lorenzo Cavallaro. 2025 · 2025
Closest in time.
Explanation as a Watermark: Towards Harmless and Multi-bit Model Ownership Verification via Watermarking Feature Attribution. In Network and Distributed System Security Symposium (NDSS)
Shuo Shao, Yiming Li, Hongwei Yao, Yiling He, Zhan Qin, and Kui Ren. 2025 · 2025
Closest in time.
Exploring Large Language Models for Semantic Analysis and Categorization of Android Malware
Brandon J Walton, Mst Eshita Khatun, James M Ghawaly, and Aisha Ali-Gombe. 2025 · 2025
Closest in time.
Apppoet: Large language model based android malware detection via multi-view prompt engineering
Wenxiang Zhao, Juntao Wu, and Zhaoyi Meng. 2025 · 2025
Closest in time.
MsDroid: Identifying Malicious Snippets for Android Malware Detection
Yiling He, Yiping Liu, Lei Wu, Ziqi Yang, Kui Ren, and Zhan Qin. 2023a · 2039
Closest in time.