Fetching the paper…
Reading the bibliography…
Decompilation aims to convert binary code to high-level source code, but traditional tools like Ghidra often produce results that are difficult to read and execute.
Omer Katz, Yuval Olshaker, Yoav Goldberg, and Eran Yahav. 2019 · 1905
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew. 2019 · 1910
Earlier work this paper cites.
Advanced compiler design and implementation
Steven S. Muchnick. 1997 · 1997
Earlier work this paper cites.
Identifying and filtering near-duplicate documents
Andrei Z Broder. 2000 · 2000
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2001
Earlier work this paper cites.
Decompiling java bytecode: Problems, traps and pitfalls
Jerome Miecznikowski and Laurie J. Hendren. 2002 · 2002
Earlier work this paper cites.
Using the gnu compiler collection
Richard M Stallman et al. 2003 · 2003
Earlier work this paper cites.
Llvm: A compilation framework for lifelong program analysis & transformation
Chris Lattner and Vikram Adve. 2004 · 2004
Earlier work this paper cites.
A new algorithm for identifying loops in decompilation
Tao Wei, Jian Mao, Wei Zou, and Yu Chen. 2007 · 2007
Earlier work this paper cites.
Decompiling android
Godfrey Nolan. 2012 · 2012
Earlier work this paper cites.
Native x86 decompilation using semantics-preserving structural analysis and iterative control-flow structuring
David Brumley, JongHyup Lee, Edward J. Schwartz, and Maverick Woo. 2013 · 2013
Earlier work this paper cites.
Obfuscator-LLVM – software protection for the masses
Pascal Junod, Julien Rinaldini, Johan Wehrli, and Julie Michielin. 2015 · 2015
Earlier work this paper cites.
Ramblr: Making reassembly great again
Ruoyu Wang, Yan Shoshitaishvili, Antonio Bianchi, Aravind Machiry, John Grosen, Paul Grosen, Christopher Kruegel, and Giovanni Vigna. 2017 · 2017
Earlier work this paper cites.
Using recurrent neural networks for decompilation
Deborah S. Katz, Jason Ruchti, and Eric M. Schulte. 2018 · 2018
Earlier work this paper cites.
DIRE: A neural approach to decompiled identifier naming
Jeremy Lacomis, Pengcheng Yin, Edward J. Schwartz, Miltiadis Allamanis, Claire Le Goues, Graham Neubig, and Bogdan Vasilescu. 2019 · 2019
Earlier work this paper cites.
Retrowrite: Statically instrumenting cots binaries for fuzzing and sanitization
Sushant Dinesh, Nathan Burow, Dongyan Xu, and Mathias Payer. 2020 · 2020
Cited alongside, same era.
How far we have come: testing decompilation correctness of c decompilers
Zhibo Liu and Shuai Wang. 2020a · 2020
Cited alongside, same era.
How far we have come: testing decompilation correctness of C decompilers
Zhibo Liu and Shuai Wang. 2020b · 2020
Cited alongside, same era.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Pondé de Oliveira Pinto, Jared Kaplan, Harrison Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Joshua Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. 2021 · 2021
Code llama: Open foundation models for code
Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, Artyom Kozhevnikov, Ivan Evtimov, Joanna Bitton, Manish Bhatt, Cristian Canton-Ferrer, Aaron Grattafiori, Wenhan Xiong, Alexandre Défossez, Jade Copet, Faisal Azhar, Hugo Touvron, Louis Martin, Nicolas Usunier, Thomas Scialom, and Gabriel Synnaeve. 2023 · 2023
Later among the works it cites.
Is ChatGPT a good NLG evaluator? a preliminary study
Jiaan Wang, Yunlong Liang, Fandong Meng, Zengkui Sun, Haoxiang Shi, Zhixu Li, Jinan Xu, Jianfeng Qu, and Jie Zhou. 2023 · 2023
Later among the works it cites.
Refining decompiled C code with large language models
Wai Kin Wong, Huaijin Wang, Zongjie Li, Zhibo Liu, Shuai Wang, Qiyi Tang, Sen Nie, and Shi Wu. 2023 · 2023
Later among the works it cites.
Lmpa: Improving decompilation by synergy of large language model and program analysis
Xiangzhe Xu, Zhuo Zhang, Shiwei Feng, Yapeng Ye, Zian Su, Nan Jiang, Siyuan Cheng, Lin Tan, and Xiangyu Zhang. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
ANGHABENCH: A suite with one million compilable C benchmarks for code-size reduction
Anderson Faustino da Silva, Bruno Conde Kind, José Wesley de Souza Magalhães, Jerônimo Nunes Rocha, Breno Campos Ferreira Guimarães, and Fernando Magno Quintão Pereira. 2021 · 2021
Cited alongside, same era.
Dobf: A deobfuscation pre-training objective for programming languages
Marie-Anne Lachaux, Baptiste Roziere, Marc Szafraniec, and Guillaume Lample. 2021 · 2021
Cited alongside, same era.
DIRECT : A transformer-based model for decompiled identifier renaming
Vikram Nitin, Anthony Saieva, Baishakhi Ray, and Gail Kaiser. 2021 · 2021
Cited alongside, same era.
Exebench: An ml-scale dataset of executable c functions
Jordi Armengol-Estapé, Jackson Woodruff, Alexander Brauckmann, José Wesley de Souza Magalhães, and Michael F. P. O’Boyle. 2022 · 2022
Cited alongside, same era.
Beyond the C: retargetable decompilation using neural machine translation
Iman Hosseini and Brendan Dolan-Gavitt. 2022 · 2022
Cited alongside, same era.
Slade: A portable small language model decompiler for optimized assembler
Jordi Armengol-Estapé, Jackson Woodruff, Chris Cummins, and Michael F. P. O’Boyle. 2023 · 2023
Cited alongside, same era.
Nova + {}^{\mbox{+}} : Generative language models for binaries
Nan Jiang, Chengxiao Wang, Kevin Liu, Xiangzhe Xu, Lin Tan, and Xiangyu Zhang. 2023 · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023 · 2023
Cited alongside, same era.
Yi-coder
01-AI. 2024 · 2024
Closest in time.
Meta large language model compiler: Foundation models of compiler optimization
Chris Cummins, Volker Seeker, Dejan Grubisic, Baptiste Roziere, Jonas Gehring, Gabriel Synnaeve, and Hugh Leather. 2024 · 2024
Closest in time.
The gnu c library
Inc. Free Software Foundation. 2024 · 2024
Closest in time.
Deepseek-coder: When the large language model meets programming–the rise of code intelligence
Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Y Wu, YK Li, et al. 2024 · 2024
Closest in time.
Ida pro: a cross-platform multi-processor disassembler and debugger
Hex-Rays. 2024 · 2024
Closest in time.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Oriol Vinyals, Jack W. Rae, and Laurent Sifre. 2024 · 2024
Closest in time.
Degpt: Optimizing decompiler output with llm
Peiwei Hu, Ruigang Liang, and Kai Chen. 2024 · 2024
Closest in time.
Machine language model
Huaqingweiyang. 2024 · 2024
Closest in time.
Codestral: Empowering developers and democratising coding with mistral ai
Mistral-AI. 2024 · 2024
Closest in time.
Leveraging generative models to recover variable names from stripped binary
Xiangzhe Xu, Zhuo Zhang, Zian Su, Ziyang Huang, Shiwei Feng, Yapeng Ye, Nan Jiang, Danning Xie, Siyuan Cheng, Lin Tan, and Xiangyu Zhang. 2024 · 2024
Closest in time.
Llamafactory: Unified efficient fine-tuning of 100+ language models
Yaowei Zheng, Richong Zhang, Junhao Zhang, Yanhan Ye, Zheyan Luo, and Yongqiang Ma. 2024 · 2024
Closest in time.