Fetching the paper…
Reading the bibliography…
Binary code summarization, while invaluable for understanding code semantics, is challenging due to its labor-intensive nature.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
Scikit-learn: Machine learning in python
Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al · 2011
Earlier work this paper cites.
srcml: An infrastructure for the exploration, analysis, and manipulation of source code: A tool demonstration
Michael L Collard, Michael John Decker, and Jonathan I Maletic · 2013
Earlier work this paper cites.
The definitive antlr 4 reference
Terence Parr · 2013
Earlier work this paper cites.
Summarizing source code using a neural attention model
Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, and Luke Zettlemoyer · 2016
Earlier work this paper cites.
Sok:(state of) the art of war: Offensive techniques in binary analysis
Yan Shoshitaishvili, Ruoyu Wang, Christopher Salls, Nick Stephens, Mario Polino, Andrew Dutcher, John Grosen, Siji Feng, Christophe Hauser, Christopher Kruegel, et al · 2016
Earlier work this paper cites.
Helping johnny to analyze malware: A usability-optimized decompiler and malware analysis user study
Khaled Yakdan, Sergej Dechand, Elmar Gerhards-Padilla, and Matthew Smith · 2016
Earlier work this paper cites.
A survey of symbolic execution techniques
Roberto Baldoni, Emilio Coppa, Daniele Cono D’elia, Camil Demetrescu, and Irene Finocchi · 2018
Earlier work this paper cites.
Vulseeker: A semantic learning based vulnerability seeker for cross-platform binary
Jian Gao, Xin Yang, Ying Fu, Yu Jiang, and Jiaguang Sun · 2018
Earlier work this paper cites.
Taku Kudo and John Richardson · 2018
Earlier work this paper cites.
Codesearchnet challenge: Evaluating the state of semantic code search
Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt · 2019
Earlier work this paper cites.
Dynamic malware analysis in the modern era—a state of the art survey
Ori Or-Meir, Nir Nissim, Yuval Elovici, and Lior Rokach · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 2019
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Neural reverse engineering of stripped binaries using augmented control flow graphs
Yaniv David, Uri Alon, and Eran Yahav · 2020
Earlier work this paper cites.
Model compression and hardware acceleration for neural networks: A comprehensive survey
Lei Deng, Guoqi Li, Song Han, Luping Shi, and Yuan Xie · 2020
Earlier work this paper cites.
Codebert: A pre-trained model for programming and natural languages
Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, et al · 2020
Earlier work this paper cites.
Graphcodebert: Pre-training code representations with data flow
Daya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng, Duyu Tang, Shujie Liu, Long Zhou, Nan Duan, Alexey Svyatkovskiy, Shengyu Fu, et al · 2020
Earlier work this paper cites.
Xda: Accurate, robust disassembly with transfer learning
Kexin Pei, Jonas Guan, David Williams-King, Junfeng Yang, and Suman Jana · 2020
Earlier work this paper cites.
Trex: Learning execution semantics from micro-traces for binary similarity
Kexin Pei, Zhou Xuan, Junfeng Yang, Suman Jana, and Baishakhi Ray · 2020
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He · 2020
Earlier work this paper cites.
Scipy 1.0: fundamental algorithms for scientific computing in python
Pauli Virtanen, Ralf Gommers, Travis E Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, et al · 2020
Earlier work this paper cites.
Order matters: Semantic-aware neural networks for binary code similarity detection
Zeping Yu, Rui Cao, Qiyi Tang, Sen Nie, Junzhou Huang, and Shi Wu · 2020
Earlier work this paper cites.
Codecmr: Cross-modal retrieval for function-level binary source code matching
Zeping Yu, Wenxin Zheng, Jiaqi Wang, Qiyi Tang, Sen Nie, and Shi Wu · 2020
Earlier work this paper cites.
Variable name recovery in decompiled binary code using constrained masked language modeling
Pratyay Banerjee, Kuntal Kumar Pal, Fish Wang, and Chitta Baral · 2021
Cited alongside, same era.
Why my code summarization model does not work: Code comment improvement with category prediction
Qiuyuan Chen, Xin Xia, Han Hu, David Lo, and Shanping Li · 2021
Cited alongside, same era.
Palmtree: Learning an assembly language model for instruction embedding
Xuezixiang Li, Yu Qu, and Heng Yin · 2021
Cited alongside, same era.
Recent advances in natural language processing via large pre-trained language models: A survey
Bonan Min, Hayley Ross, Elior Sulem, Amir Pouran Ben Veyseh, Thien Huu Nguyen, Oscar Sainz, Eneko Agirre, Ilana Heintz, and Dan Roth · 2021
Cited alongside, same era.
Deep learning–based text classification: a comprehensive review
Shervin Minaee, Nal Kalchbrenner, Erik Cambria, Narjes Nikzad, Meysam Chenaghlu, and Jianfeng Gao · 2021
Cited alongside, same era.
https://blog.virustotal.com/2023/04/introducing-virustotal-code-insight.html
Introducing virustotal code insight: Empowering threat analysis with generative ai · 2023
Closest in time.
https://huggingface.co/meta-llama
Meta llama 2 · 2023
Closest in time.
https://www.hex-rays.com/products/ida/support/idapython_docs/ida_hexrays.html
Module ida_hexrays · 2023
Closest in time.
https://www.hex-rays.com/products/ida/support/idapython_docs/ida_idaapi.html
Module ida_idaapi · 2023
Closest in time.
https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderboard
Open llm leaderboard · 2023
Closest in time.
https://platform.openai.com/docs/api-reference/chat/create
Openai api reference - create chat completion · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stateformer: Fine-grained type recovery from binaries using generative state modeling
Kexin Pei, Jonas Guan, Matthew Broughton, Zhongtian Chen, Songchen Yao, David Williams-King, Vikas Ummadisetty, Junfeng Yang, Baishakhi Ray, and Suman Jana · 2021
Cited alongside, same era.
Yue Wang, Weishi Wang, Shafiq Joty, and Steven CH Hoi · 2021
Cited alongside, same era.
Augmenting decompiler output with learned variable names and types
Qibin Chen, Jeremy Lacomis, Edward J Schwartz, Claire Le Goues, Graham Neubig, and Bogdan Vasilescu · 2022
Cited alongside, same era.
Large language models can self-improve
Jiaxin Huang, Shixiang Shane Gu, Le Hou, Yuexin Wu, Xuezhi Wang, Hongkun Yu, and Jiawei Han · 2022
Cited alongside, same era.
Symlm: Predicting function names in stripped binaries via context-sensitive execution-aware code embeddings
Xin Jin, Kexin Pei, Jun Yeon Won, and Zhiqiang Lin · 2022
Cited alongside, same era.
The stack: 3 tb of permissively licensed source code
Denis Kocetkov, Raymond Li, Loubna Ben Allal, Jia Li, Chenghao Mou, Carlos Muñoz Ferrandis, Yacine Jernite, Margaret Mitchell, Sean Hughes, Thomas Wolf, et al · 2022
Cited alongside, same era.
Neudep: neural binary memory dependence analysis
Kexin Pei, Dongdong She, Michael Wang, Scott Geng, Zhou Xuan, Yaniv David, Junfeng Yang, Suman Jana, and Baishakhi Ray · 2022
Cited alongside, same era.
https://ghidra.re/ghidra_docs/api/ghidra/app/decompiler/package-summary.html
Package ghidra.app.decompiler · 2023
Closest in time.
https://github.com/eliben/pyelftools
pyelftools - parsing elf and dwarf in python · 2023
Closest in time.
https://www.sbert.net/docs/pretrained_models.html#pretrained-models
Sentence-transformers - pretrained models · 2023
Closest in time.
https://blog.virustotal.com/2023/05/vt-code-insight-updates-and-q-on.html
Vt code insight: Updates and q&a on purpose, challenges, and evolution · 2023
Closest in time.
https://ghidra-sre.org/ , 2023
Ghidra · 2023
Closest in time.
https://platform.openai.com/docs/guides/rate-limits/overview , Accessed: 2023-09-29
Openai - rate limits · 2023
Closest in time.
Extending source code pre-trained language models to summarise decompiled binarie
Ali Al-Kaswan, Toufique Ahmed, Maliheh Izadi, Anand Ashok Sawant, Premkumar Devanbu, and Arie van Deursen · 2023
Closest in time.
Syzdescribe: Principled, automated, static generation of syscall descriptions for kernel drivers
Yu Hao, Guoren Li, Xiaochen Zou, Weiteng Chen, Shitong Zhu, Zhiyun Qian, and Ardalan Amiri Sani · 2023
Closest in time.
Guiding large language models via directional stimulus prompting
Zekun Li, Baolin Peng, Pengcheng He, Michel Galley, Jianfeng Gao, and Xifeng Yan · 2023
Closest in time.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig · 2023
Closest in time.
Vulhawk: Cross-architecture vulnerability detection with entropy-based binary code search
Zhenhao Luo, Pengfei Wang, Baosheng Wang, Yong Tang, Wei Xie, Xu Zhou, Danjun Liu, and Kai Lu · 2023
Closest in time.
Gpt-4 technical report
R OpenAI · 2023
Closest in time.
Automated generation of security-centric descriptions for smart contract bytecode
Yu Pan, Zhichao Xu, Levi Taiji Li, Yunhe Yang, and Mu Zhang · 2023
Closest in time.
Xfl: Naming functions in binaries with extreme multi-label learning
James Patrick-Evans, Moritz Dannehl, and Johannes Kinder · 2023
Closest in time.
Automatic prompt optimization with" gradient descent" and beam search
Reid Pryzant, Dan Iter, Jerry Li, Yin Tat Lee, Chenguang Zhu, and Michael Zeng · 2023
Closest in time.
Code llama: Open foundation models for code
Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, et al · 2023
Closest in time.
Gpt-4 might just be a bloated, pointless mess - the atlantic
Jacob Stern · 2023
Closest in time.
Is chatgpt the ultimate programming assistant–how far is it?
Haoye Tian, Weiqi Lu, Tsz On Li, Xunzhu Tang, Shing-Chi Cheung, Jacques Klein, and Tegawendé F Bissyandé · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Closest in time.
Large language models as optimizers
Chengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu, Quoc V Le, Denny Zhou, and Xinyun Chen · 2023
Closest in time.
Exploring the limits of chatgpt for query or aspect-based text summarization
Xianjun Yang, Yan Li, Xinlu Zhang, Haifeng Chen, and Wei Cheng · 2023
Closest in time.