Fetching the paper…
Reading the bibliography…
Transformer based code models have impressive performance in many software engineering tasks.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Two sides of the same coin: Exploiting the impact of identifiers in neural code comprehension. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE) . IEEE, 1933–1945
Shuzheng Gao, Cuiyun Gao, Chaozheng Wang, Jun Sun, David Lo, and Yue Yu. 2023 · 1945
Earlier work this paper cites.
A program data flow analysis procedure
Frances E. Allen and John Cocke. 1976 · 1976
Earlier work this paper cites.
The program dependence graph and its use in optimization
Jeanne Ferrante, Karl J Ottenstein, and Joe D Warren. 1987 · 1987
Earlier work this paper cites.
Types and programming languages
Benjamin C Pierce. 2002 · 2002
Earlier work this paper cites.
Automatic reverse engineering of data structures from binary execution. In Proceedings of the 11th Annual Information Security Symposium . 1–1
Zhiqiang Lin, Xiangyu Zhang, and Dongyan Xu. 2010 · 2010
Earlier work this paper cites.
Mining multi-label data
Grigorios Tsoumakas, Ioannis Katakis, and Ioannis Vlahavas. 2010 · 2010
Earlier work this paper cites.
Malware Images: Visualization and Automatic Classification. In Proceedings of the 8th International Symposium on Visualization for Cyber Security (Pittsburgh, Pennsylvania, USA) (VizSec ’11) . Association for Computing Machinery, New York, NY, USA, Article 4, 7 pages
L. Nataraj, S. Karthikeyan, G. Jacob, and B. S. Manjunath. 2011 · 2011
Earlier work this paper cites.
Anatomization and protection of mobile apps’ location privacy threats. In 24th USENIX Security Symposium (USENIX Security 15) . 753–768
Kassem Fawaz, Huan Feng, and Kang G Shin. 2015 · 2015
Earlier work this paper cites.
Effective Approaches to Attention-based Neural Machine Translation. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing . 1412–1421
Minh-Thang Luong, Hieu Pham, and Christopher D Manning. 2015 · 2015
Earlier work this paper cites.
Learning to represent programs with graphs
Miltiadis Allamanis, Marc Brockschmidt, and Mahmoud Khademi. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2017 · 2017
Earlier work this paper cites.
Semantics-Based Obfuscation-Resilient Binary Code Similarity Comparison with Applications to Software and Algorithm Plagiarism Detection
Lannan Luo, Jiang Ming, Dinghao Wu, Peng Liu, and Sencun Zhu. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
In-Memory Fuzzing for Binary Code Similarity Analysis. In Proceedings of the 32nd IEEE/ACM International Conference on Automated Software Engineering (Urbana-Champaign, IL, USA) (ASE 2017) . IEEE Press, 319–330
Shuai Wang and Dinghao Wu. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Deeper insights into graph convolutional networks for semi-supervised learning. In Proceedings of the AAAI conference on artificial intelligence , Vol. 32
Qimai Li, Zhichao Han, and Xiao-Ming Wu. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Graph matching networks for learning the similarity of graph structured objects. In International conference on machine learning . PMLR, 3835–3845
Yujia Li, Chenjie Gu, Thomas Dullien, Oriol Vinyals, and Pushmeet Kohli. 2019 · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Earlier work this paper cites.
BDA: practical dependence analysis for binary executables by unbiased whole-program path sampling and per-path abstract interpretation
Zhuo Zhang, Wei You, Guanhong Tao, Guannan Wei, Yonghwi Kwon, and Xiangyu Zhang. 2019 · 2019
Earlier work this paper cites.
On the bottleneck of graph neural networks and its practical implications
Uri Alon and Eran Yahav. 2020 · 2020
Earlier work this paper cites.
Learning to execute programs with instruction pointer attention graph neural networks
David Bieber, Charles Sutton, Hugo Larochelle, and Daniel Tarlow. 2020 · 2020
Earlier work this paper cites.
Measuring and relieving the over-smoothing problem for graph neural networks from the topological view. In Proceedings of the AAAI conference on artificial intelligence , Vol. 34. 3438–3445
Deli Chen, Yankai Lin, Wei Li, Peng Li, Jie Zhou, and Xu Sun. 2020 · 2020
Earlier work this paper cites.
Deepbindiff: Learning program-wide code representations for binary diffing. In Network and distributed system security symposium
Yue Duan, Xuezixiang Li, Jinghan Wang, and Heng Yin. 2020 · 2020
Earlier work this paper cites.
CodeBERT: A Pre-Trained Model for Programming and Natural Languages. In Findings of the Association for Computational Linguistics: EMNLP 2020 . 1536–1547
Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, et al · 2020
Earlier work this paper cites.
Graphcodebert: Pre-training code representations with data flow
Daya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng, Duyu Tang, Shujie Liu, Long Zhou, Nan Duan, Alexey Svyatkovskiy, Shengyu Fu, et al · 2020
Earlier work this paper cites.
GRAPHSPY: Fused Program Semantic-Level Embedding via Graph Neural Networks for Dead Store Detection
Yixin Guo, Pengcheng Li, Yingwei Luo, Xiaolin Wang, and Zhenlin Wang. 2020a · 2020
Cited alongside, same era.
Learning and evaluating contextual embedding of source code. In International conference on machine learning . PMLR, 5110–5121
Aditya Kanade, Petros Maniatis, Gogul Balakrishnan, and Kensen Shi. 2020 · 2020
Cited alongside, same era.
Trex: Learning execution semantics from micro-traces for binary similarity
Kexin Pei, Zhou Xuan, Junfeng Yang, Suman Jana, and Baishakhi Ray. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Cited alongside, same era.
Graphcode2vec: Generic code embedding via lexical and program dependence analyses. In Proceedings of the 19th International Conference on Mining Software Repositories . 524–536
Wei Ma, Mengjie Zhao, Ezekiel Soremekun, Qiang Hu, Jie M Zhang, Mike Papadakis, Maxime Cordy, Xiaofei Xie, and Yves Le Traon. 2022 · 2022
Later among the works it cites.
How to better utilize code graphs in semantic code search?. In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering . 722–733
Yucen Shi, Ying Yin, Zhengkui Wang, David Lo, Tao Zhang, Xin Xia, Yuhai Zhao, and Bowen Xu. 2022 · 2022
Later among the works it cites.
A systematic evaluation of large language models of code. In Proceedings of the 6th ACM SIGPLAN International Symposium on Machine Programming . 1–10
Frank F Xu, Uri Alon, Graham Neubig, and Vincent Josua Hellendoorn. 2022 · 2022
Later among the works it cites.
DeepDi: Learning a Relational Graph Convolutional Network Model on Instructions for Fast and Accurate Disassembly. In 31st USENIX Security Symposium (USENIX Security 22) . 2709–2725
Sheng Yu, Yu Qu, Xunchao Hu, and Heng Yin. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zeyu Sun, Qihao Zhu, Yingfei Xiong, Yican Sun, Lili Mou, and Lu Zhang. 2020 · 2020
Cited alongside, same era.
Beyond homophily in graph neural networks: Current limitations and effective designs
Jiong Zhu, Yujun Yan, Lingxiao Zhao, Mark Heimann, Leman Akoglu, and Danai Koutra. 2020 · 2020
Cited alongside, same era.
Unified Pre-training for Program Understanding and Generation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . 2655–2668
Wasi Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2021 · 2021
Cited alongside, same era.
Infercode: Self-supervised learning of code representations by predicting subtrees. In 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE) . IEEE, 1186–1197
Nghi DQ Bui, Yijun Yu, and Lingxiao Jiang. 2021a · 2021
Cited alongside, same era.
PLUR: A unifying, graph-based view of program learning, understanding, and repair
Zimin Chen, Vincent J Hellendoorn, Pascal Lamblin, Petros Maniatis, Pierre-Antoine Manzagol, Daniel Tarlow, and Subhodeep Moitra. 2021 · 2021
Cited alongside, same era.
Identifying authorship style in malicious binaries: techniques, challenges & datasets
Jason Gray, Daniele Sgandurra, and Lorenzo Cavallaro. 2021 · 2021
Cited alongside, same era.
Codexglue: A machine learning benchmark dataset for code understanding and generation
Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin Clement, Dawn Drain, Daxin Jiang, Duyu Tang, et al · 2021
Cited alongside, same era.
A comprehensive study on learning-based PE malware family classification methods. In Proceedings of the 29th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering . 1314–1325
Yixuan Ma, Shuang Liu, Jiajun Jiang, Guanhong Chen, and Keqiu Li. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Later among the works it cites.
A study on prompt design, advantages and limitations of chatgpt for deep learning program repair
Jialun Cao, Meiziniu Li, Ming Wen, and Shing-chi Cheung. 2023 · 2023
Later among the works it cites.
TRACED: Execution-aware Pre-training for Source Code
Yangruibo Ding, Ben Steenhoek, Kexin Pei, Gail Kaiser, Wei Le, and Baishakhi Ray. 2023b · 2023
Later among the works it cites.
Towards learning generalizable code embeddings using task-agnostic graph convolutional networks
Zishuo Ding, Heng Li, Weiyi Shang, and Tse-Hsun Chen. 2023a · 2023
Later among the works it cites.
Automated repair of programs from large language models. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE) . IEEE, 1469–1481
Zhiyu Fan, Xiang Gao, Martin Mirchev, Abhik Roychoudhury, and Shin Hwei Tan. 2023 · 2023
Later among the works it cites.
RepresentThemAll: A Universal Learning Representation of Bug Reports. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE) . IEEE, 602–614
Sen Fang, Tao Zhang, Youshuai Tan, He Jiang, Xin Xia, and Xiaobing Sun. 2023a · 2023
Later among the works it cites.
A study on the impact of pre-trained model on Just-In-Time defect prediction
Yuxiang Guo, Xiaopeng Gao, Zhenyu Zhang, WK Chan, and Bo Jiang. 2023 · 2023
Later among the works it cites.
A powerful disassembler and a versatile debugger
IDA Pro 2023 · 2023
Later among the works it cites.
CCT5: A Code-Change-Oriented Pre-Trained Model
Bo Lin, Shangwen Wang, Zhongxin Liu, Yepang Liu, Xin Xia, and Xiaoguang Mao. 2023 · 2023
Later among the works it cites.
An empirical study of text-based machine learning models for vulnerability detection
Kollin Napier, Tanmay Bhowmik, and Shaowei Wang. 2023 · 2023
Later among the works it cites.
Domain Knowledge Matters: Improving Prompts with Fix Templates for Repairing Python Type Errors
Yun Peng, Shuzheng Gao, Cuiyun Gao, Yintong Huo, and Michael R Lyu. 2023 · 2023
Later among the works it cites.
Code llama: Open foundation models for code
Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, et al · 2023
Later among the works it cites.
Model-Agnostic Syntactical Information for Pre-Trained Programming Language Models
Iman Saberi and Fatemeh H Fard. 2023 · 2023
Later among the works it cites.
SEAL: Integrating Program Analysis and Repository Mining
Florian Sattler, Sebastian Böhm, Philipp Dominik Schubert, Norbert Siegmund, and Sven Apel. 2023 · 2023
Later among the works it cites.
Towards Efficient Fine-tuning of Pre-trained Code Models: An Experimental Study and Beyond
Ensheng Shi, Yanlin Wang, Hongyu Zhang, Lun Du, Shi Han, Dongmei Zhang, and Hongbin Sun. 2023 · 2023
Later among the works it cites.
How Effective Are Neural Networks for Fixing Security Vulnerabilities
Yi Wu, Nan Jiang, Hung Viet Pham, Thibaud Lutellier, Jordan Davis, Lin Tan, Petr Babkin, and Sameena Shah. 2023 · 2023
Later among the works it cites.
Improving Binary Code Similarity Transformer Models by Semantics-Driven Instruction Deemphasis. In Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis (Seattle, WA, USA) (ISSTA 2023) . Association for Computing Machinery, New York, NY, USA, 1106–1118
Xiangzhe Xu, Shiwei Feng, Yapeng Ye, Guangyu Shen, Zian Su, Siyuan Cheng, Guanhong Tao, Qingkai Shi, Zhuo Zhang, and Xiangyu Zhang. 2023a · 2023
Later among the works it cites.
Xiangzhe Xu, Zhou Xuan, Shiwei Feng, Siyuan Cheng, Yapeng Ye, Qingkai Shi, Guanhong Tao, Le Yu, Zhuo Zhang, and Xiangyu Zhang. 2023b · 2023
Later among the works it cites.
kTrans: Knowledge-Aware Transformer for Binary Code Embedding
Wenyu Zhu, Hao Wang, Yuchen Zhou, Jiaming Wang, Zihan Sha, Zeyu Gao, and Chao Zhang. 2023 · 2023
Later among the works it cites.
SIGMADIFF: Semantics-Aware Deep Graph Matching for Pseudocode Diffing. In Network and Distributed System Security (NDSS) Symposium 2024
Lian Gao, Yu Qu, Sheng Yu, Yue Duan, and Heng Yin. [n. d.] · 2024
Closest in time.
Dataflow Analysis-Inspired Deep Learning for Efficient Vulnerability Detection. In Proceedings of the 46th IEEE/ACM International Conference on Software Engineering . 1–13
Benjamin Steenhoek, Hongyang Gao, and Wei Le. 2024 · 2024
Closest in time.
How machine learning is solving the binary function similarity problem. In 31st USENIX Security Symposium (USENIX Security 22) . 2099–2116
Andrea Marcelli, Mariano Graziano, Xabier Ugarte-Pedrero, Yanick Fratantonio, Mohamad Mansouri, and Davide Balzarotti. 2022 · 2099
Closest in time.