Fetching the paper…
Reading the bibliography…
Recognizing vulnerabilities in stripped binary files presents a significant challenge in software security.
Long Short-term Memory
S Hochreiter. 1997 · 1997
Earlier work this paper cites.
TIE: Principled reverse engineering of types in binary programs
JongHyup Lee, Thanassis Avgerinos, and David Brumley. 2011 · 2011
Earlier work this paper cites.
Polymorphic type inference for machine code. In Proceedings of the 37th ACM SIGPLAN Conference on Programming Language Design and Implementation . 27–41
Matt Noonan, Alexey Loginov, and David Cok. 2016 · 2016
Earlier work this paper cites.
Sok:(state of) the art of war: Offensive techniques in binary analysis. In 2016 IEEE symposium on security and privacy (SP) . IEEE, 138–157
Yan Shoshitaishvili, Ruoyu Wang, Christopher Salls, Nick Stephens, Mario Polino, Andrew Dutcher, John Grosen, Siji Feng, Christophe Hauser, Christopher Kruegel, et al · 2016
Earlier work this paper cites.
Attention is all you need
A Vaswani. 2017 · 2017
Earlier work this paper cites.
Angr-the next generation of binary analysis. In 2017 IEEE Cybersecurity Development (SecDev) . IEEE, 8–9
Fish Wang and Yan Shoshitaishvili. 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin. 2018 · 2018
Earlier work this paper cites.
Vulseeker: A semantic learning based vulnerability seeker for cross-platform binary. In Proceedings of the 33rd ACM/IEEE International Conference on Automated Software Engineering . 896–899
Jian Gao, Xin Yang, Ying Fu, Yu Jiang, and Jiaguang Sun. 2018 · 2018
Earlier work this paper cites.
Debin: Predicting debug information in stripped binaries. In Proceedings of the 2018 ACM SIGSAC Conference on Computer and Communications Security . 1667–1680
Jingxuan He, Pesho Ivanov, Petar Tsankov, Veselin Raychev, and Martin Vechev. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford. 2018 · 2018
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu. 2019 · 2019
Earlier work this paper cites.
Devign: Effective vulnerability identification by learning comprehensive program semantics via graph neural networks
Yaqin Zhou, Shangqing Liu, Jingkai Siow, Xiaoning Du, and Yang Liu. 2019 · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown. 2020 · 2020
Earlier work this paper cites.
Cati: Context-assisted type inference from stripped binaries. In 2020 50th Annual IEEE/IFIP International Conference on Dependable Systems and Networks (DSN) . IEEE, 88–98
Ligeng Chen, Zhongling He, and Bing Mao. 2020 · 2020
Earlier work this paper cites.
Neural reverse engineering of stripped binaries using augmented control flow graphs
Yaniv David, Uri Alon, and Eran Yahav. 2020 · 2020
Earlier work this paper cites.
AC/C++ code vulnerability dataset with code changes and CVE summaries. In Proceedings of the 17th International Conference on Mining Software Repositories . 508–512
Jiahao Fan, Yi Li, Shaohua Wang, and Tien N Nguyen. 2020 · 2020
Earlier work this paper cites.
Codebert: A pre-trained model for programming and natural languages
Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, et al · 2020
Earlier work this paper cites.
Pmp: Cost-effective forced execution with probabilistic memory pre-planning. In 2020 IEEE Symposium on Security and Privacy (SP) . IEEE, 1121–1138
Wei You, Zhuo Zhang, Yonghwi Kwon, Yousra Aafer, Fei Peng, Yu Shi, Carson Harmon, and Xiangyu Zhang. 2020 · 2020
Earlier work this paper cites.
Graph-bert: Only attention is needed for learning graph representations
Jiawei Zhang, Haopeng Zhang, Congying Xia, and Li Sun. 2020 · 2020
Earlier work this paper cites.
Unified pre-training for program understanding and generation
Wasi Uddin Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. 2021 · 2021
Earlier work this paper cites.
Deep learning based vulnerability detection: Are we there yet?
Saikat Chakraborty, Rahul Krishna, Yangruibo Ding, and Baishakhi Ray. 2021 · 2021
Earlier work this paper cites.
{ \{ SelectiveTaint } \} : Efficient Data Flow Tracking With Static Binary Rewriting. In 30th USENIX Security Symposium (USENIX Security 21) . 1665–1682
Sanchuan Chen, Zhiqiang Lin, and Yinqian Zhang. 2021 · 2021
Earlier work this paper cites.
Reverse engineering learned optimizers reveals known and novel mechanisms
Niru Maheswaranathan, David Sussillo, Luke Metz, Ruoxi Sun, and Jascha Sohl-Dickstein. 2021 · 2021
Earlier work this paper cites.
Direct: A transformer-based model for decompiled identifier renaming. In Proceedings of the 1st Workshop on Natural Language Processing for Programming (NLP4Prog 2021) . 48–57
Vikram Nitin, Anthony Saieva, Baishakhi Ray, and Gail Kaiser. 2021 · 2021
Earlier work this paper cites.
Sok: All you ever wanted to know about x86/x64 binary disassembly but were afraid to ask. In 2021 IEEE symposium on security and privacy (SP) . IEEE, 833–851
Chengbin Pang, Ruotong Yu, Yaohui Chen, Eric Koskinen, Georgios Portokalidis, Bing Mao, and Jun Xu. 2021 · 2021
Earlier work this paper cites.
Yue Wang, Weishi Wang, Shafiq Joty, and Steven CH Hoi. 2021 · 2021
Earlier work this paper cites.
Osprey: Recovery of variable and data structure via probabilistic analysis for stripped binary. In 2021 IEEE Symposium on Security and Privacy (SP) . IEEE, 813–832
Zhuo Zhang, Yapeng Ye, Wei You, Guanhong Tao, Wen-chuan Lee, Yonghwi Kwon, Yousra Aafer, and Xiangyu Zhang. 2021 · 2021
Earlier work this paper cites.
Augmenting decompiler output with learned variable names and types. In 31st USENIX Security Symposium (USENIX Security 22) . 4327–4343
Qibin Chen, Jeremy Lacomis, Edward J Schwartz, Claire Le Goues, Graham Neubig, and Bogdan Vasilescu. 2022 · 2022
Earlier work this paper cites.
Investigating graph embedding methods for cross-platform binary code similarity detection. In 2022 IEEE 7th European Symposium on Security and Privacy (EuroS&P) . IEEE, 60–73
Victor Cochard, Damian Pfammatter, Chi Thang Duong, and Mathias Humbert. 2022 · 2022
Earlier work this paper cites.
A survey on in-context learning
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Tianyu Liu, et al · 2022
Earlier work this paper cites.
Incoder: A generative model for code infilling and synthesis
Daniel Fried, Armen Aghajanyan, Jessy Lin, Sida Wang, Eric Wallace, Freda Shi, Ruiqi Zhong, Wen-tau Yih, Luke Zettlemoyer, and Mike Lewis. 2022 · 2022
Earlier work this paper cites.
Linevul: A transformer-based line-level vulnerability prediction. In Proceedings of the 19th International Conference on Mining Software Repositories . 608–620
Michael Fu and Chakkrit Tantithamthavorn. 2022 · 2022
Earlier work this paper cites.
Risotto: a dynamic binary translator for weak memory model architectures. In Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1 . 107–122
Redha Gouicem, Dennis Sprokholt, Jasper Ruehl, Rodrigo CO Rocha, Tom Spink, Soham Chakraborty, and Pramod Bhatotia. 2022 · 2022
Earlier work this paper cites.
Vulberta: Simplified source code pre-training for vulnerability detection. In 2022 International joint conference on neural networks (IJCNN) . IEEE, 1–8
Hazim Hanif and Sergio Maffeis. 2022 · 2022
Earlier work this paper cites.
Revisiting binary code similarity analysis using interpretable feature engineering and lessons learned
Dongkwan Kim, Eunsoo Kim, Sang Kil Cha, Sooel Son, and Yongdae Kim. 2022 · 2022
Earlier work this paper cites.
Text and code embeddings by contrastive pre-training
Arvind Neelakantan, Tao Xu, Raul Puri, Alec Radford, Jesse Michael Han, Jerry Tworek, Qiming Yuan, Nikolas Tezak, Jong Wook Kim, Chris Hallacy, et al · 2022
Earlier work this paper cites.
Codegen: An open large language model for code with multi-turn program synthesis
Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. 2022 · 2022
Cited alongside, same era.
Learning approximate execution semantics from traces for binary function similarity
Kexin Pei, Zhou Xuan, Junfeng Yang, Suman Jana, and Baishakhi Ray. 2022b · 2022
Cited alongside, same era.
Arbiter: Bridging the static and dynamic divide in vulnerability discovery on binary programs. In 31st USENIX Security Symposium (USENIX Security 22) . 413–430
Jayakrishna Vadayath, Moritz Eckert, Kyle Zeng, Nicolaas Weideman, Gokulkrishna Praveen Menon, Yanick Fratantonio, Davide Balzarotti, Adam Doupé, Tiffany Bao, Ruoyu Wang, et al · 2022
Cited alongside, same era.
Jtrans: Jump-aware transformer for binary code similarity detection. In Proceedings of the 31st ACM SIGSOFT International Symposium on Software Testing and Analysis . 1–13
Hao Wang, Wenjie Qu, Gilad Katz, Wenyu Zhu, Zeyu Gao, Han Qiu, Jianwei Zhuge, and Chao Zhang. 2022 · 2022
Cited alongside, same era.
Can a Deep Learning Model for One Architecture Be Used for Others? { \{ Retargeted-Architecture } \} Binary Code Analysis. In 32nd USENIX Security Symposium (USENIX Security 23) . 7339–7356
Junzhe Wang, Matthew Sharp, Chuxiong Wu, Qiang Zeng, and Lannan Luo. 2023 · 2023
Later among the works it cites.
How effective are neural networks for fixing security vulnerabilities. In Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis . 1282–1294
Yi Wu, Nan Jiang, Hung Viet Pham, Thibaud Lutellier, Jordan Davis, Lin Tan, Petr Babkin, and Sameena Shah. 2023 · 2023
Later among the works it cites.
Automated program repair in the era of large pre-trained language models. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE) . IEEE, 1482–1494
Chunqiu Steven Xia, Yuxiang Wei, and Lingming Zhang. 2023 · 2023
Later among the works it cites.
Impact of large language models on generating software specifications
Danning Xie, Byungwoo Yoo, Nan Jiang, Mijung Kim, Lin Tan, Xiangyu Zhang, and Judy S Lee. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Cited alongside, same era.
CodeLlaMA by Meta AI
2023 · 2023
Cited alongside, same era.
GPT-4 by OpenAI
2023 · 2023
Cited alongside, same era.
Unsupervised Binary Code Translation with Application to Code Clone Detection and Vulnerability Discovery. In Findings of the Association for Computational Linguistics: EMNLP 2023 . 14581–14592
Iftakhar Ahmad and Lannan Luo. 2023 · 2023
Cited alongside, same era.
{ \{ FirmSolo } \} : Enabling dynamic analysis of binary Linux-based { \{ IoT } \} kernel modules. In 32nd USENIX Security Symposium (USENIX Security 23) . 5021–5038
Ioannis Angelakopoulos, Gianluca Stringhini, and Manuel Egele. 2023 · 2023
Cited alongside, same era.
Teaching large language models to self-debug
Xinyun Chen, Maxwell Lin, Nathanael Schärli, and Denny Zhou. 2023 · 2023
Cited alongside, same era.
Large language models are zero-shot fuzzers: Fuzzing deep-learning libraries via large language models. In Proceedings of the 32nd ACM SIGSOFT international symposium on software testing and analysis . 423–435
Yinlin Deng, Chunqiu Steven Xia, Haoran Peng, Chenyuan Yang, and Lingming Zhang. 2023 · 2023
Cited alongside, same era.
An empirical study on using large language models for multi-intent comment generation
Mingyang Geng, Shangwen Wang, Dezun Dong, Haotian Wang, Ge Li, Zhi Jin, Xiaoguang Mao, and Xiangke Liao. 2023 · 2023
Cited alongside, same era.
Improving binary code similarity transformer models by semantics-driven instruction deemphasis. In Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis . 1106–1118
Xiangzhe Xu, Shiwei Feng, Yapeng Ye, Guangyu Shen, Zian Su, Siyuan Cheng, Guanhong Tao, Qingkai Shi, Zhuo Zhang, and Xiangyu Zhang. 2023 · 2023
Later among the works it cites.
Asteria-Pro: Enhancing Deep Learning-based Binary Code Similarity Detection by Incorporating Domain Knowledge
Shouguo Yang, Chaopeng Dong, Yang Xiao, Yiran Cheng, Zhiqiang Shi, Zhi Li, and Limin Sun. 2023 · 2023
Later among the works it cites.
{ \{ UVSCAN } \} : Detecting { \{ Third-Party } \} Component Usage Violations in { \{ IoT } \} Firmware. In 32nd USENIX Security Symposium (USENIX Security 23) . 3421–3438
Binbin Zhao, Shouling Ji, Xuhong Zhang, Yuan Tian, Qinying Wang, Yuwen Pu, Chenyang Lyu, and Raheem Beyah. 2023 · 2023
Later among the works it cites.
Big-Vul Dataset
2024 · 2024
Later among the works it cites.
Devign Dataset
2024 · 2024
Later among the works it cites.
IDA Pro
2024 · 2024
Later among the works it cites.
Juliet Test Suite for C/C++ and Java
2024 · 2024
Later among the works it cites.
Microsoft Vulnerability Dataset (MVD)
2024 · 2024
Later among the works it cites.
REVEAL Dataset
2024 · 2024
Later among the works it cites.
Software Assurance Reference Dataset (SARD)
2024 · 2024
Later among the works it cites.
Vul4J Dataset
2024 · 2024
Later among the works it cites.
Tri Dao and Albert Gu. 2024 · 2024
Later among the works it cites.
Vul-RAG: Enhancing LLM-based Vulnerability Detection via Knowledge-level RAG
Xueying Du, Geng Zheng, Kaixin Wang, Jiayi Feng, Wentai Deng, Mingwei Liu, Bihuan Chen, Xin Peng, Tao Ma, and Yiling Lou. 2024 · 2024
Later among the works it cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Later among the works it cites.
DeepSeek-Coder: When the Large Language Model Meets Programming–The Rise of Code Intelligence
Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Yu Wu, YK Li, et al · 2024
Later among the works it cites.
Code is not natural language: Unlock the power of semantics-oriented graph representation for binary code similarity detection. In 33rd USENIX Security Symposium (USENIX Security 24), PHILADELPHIA, PA
Haojie He, Xingwei Lin, Ziang Weng, Ruijie Zhao, Shuitao Gan, Libo Chen, Yuede Ji, Jiashui Wang, and Zhi Xue. 2024 · 2024
Later among the works it cites.
DeGPT: Optimizing Decompiler Output with LLM. In Proceedings 2024 Network and Distributed System Security Symposium (2024). https://api. semanticscholar. org/CorpusID , Vol. 267622140
Peiwei Hu, Ruigang Liang, and Kai Chen. 2024 · 2024
Later among the works it cites.
Pre-training by Predicting Program Dependencies for Vulnerability Analysis Tasks. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering . 1–13
Zhongxin Liu, Zhijie Tang, Junwei Zhang, Xin Xia, and Xiaohu Yang. 2024 · 2024
Later among the works it cites.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al · 2024
Later among the works it cites.
Common Vulnerabilities and Exposures
MITRE. 2024 · 2024
Later among the works it cites.
The website of cwe-416
MITRE. 2024 · 2024
Later among the works it cites.
Using an llm to help with code understanding. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering . 1–13
Daye Nam, Andrew Macvean, Vincent Hellendoorn, Bogdan Vasilescu, and Brad Myers. 2024 · 2024
Later among the works it cites.
Domain knowledge matters: Improving prompts with fix templates for repairing python type errors. In Proceedings of the 46th IEEE/ACM International Conference on Software Engineering . 1–13
Yun Peng, Shuzheng Gao, Cuiyun Gao, Yintong Huo, and Michael Lyu. 2024 · 2024
Later among the works it cites.
Dataflow analysis-inspired deep learning for efficient vulnerability detection. In Proceedings of the 46th IEEE/ACM International Conference on Software Engineering . 1–13
Benjamin Steenhoek, Hongyang Gao, and Wei Le. 2024 · 2024
Later among the works it cites.
Source Code Foundation Models are Transferable Binary Analysis Knowledge Bases
Zian Su, Xiangzhe Xu, Ziyang Huang, Kaiyuan Zhang, and Xiangyu Zhang. 2024a · 2024
Later among the works it cites.
Codeart: Better code models by attention regularization when symbols are lacking
Zian Su, Xiangzhe Xu, Ziyang Huang, Zhuo Zhang, Yapeng Ye, Jianjun Huang, and Xiangyu Zhang. 2024b · 2024
Later among the works it cites.
LLM4Decompile: Decompiling Binary Code with Large Language Models
Hanzhuo Tan, Qi Luo, Jing Li, and Yuqun Zhang. 2024 · 2024
Later among the works it cites.
LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks. In IEEE Symposium on Security and Privacy
Saad Ullah, Mingji Han, Saurabh Pujar, Hammond Pearce, Ayse Coskun, and Gianluca Stringhini. 2024 · 2024
Later among the works it cites.
ReSym: Harnessing LLMs to Recover Variable and Data Structure Symbols from Stripped Binaries
Danning Xie, Zhuo Zhang, Nan Jiang, Xiangzhe Xu, Lin Tan, and Xiangyu Zhang. 2024 · 2024
Later among the works it cites.
{ \{ LLM-Fuzzer } \} : Scaling Assessment of Large Language Model Jailbreaks. In 33rd USENIX Security Symposium (USENIX Security 24) . 4657–4674
Jiahao Yu, Xingwei Lin, Zheng Yu, and Xinyu Xing. 2024 · 2024
Later among the works it cites.
Large Language Model for Vulnerability Detection and Repair: Literature Review and Roadmap
Xin Zhou, Sicong Cao, Xiaobing Sun, and David Lo. 2024 · 2024
Later among the works it cites.