SoK: Are LLMs Reliable at Source Code Recovery? A Taxonomy and Empirical Evaluation
summary
The gist
The gist The work presents the first Systematization of Knowledge (SoK) focused specifically on LLM-assisted binary-to-source recovery, providing a granular design-centric taxonomy and systematic
In short
The work introduces a Systematization of Knowledge (SoK) framework to systematically evaluate how Large Language Models (LLMs) can recover source code from binary files. By defining four research questions and six specific metrics, the study compares various state-of-the-art decompilation methods across different architectures and languages. Findings show that iterative, multi-role agents like Agent4Decompile perform best for byte-level accuracy.
Key concepts
- Systematization of Knowledge (SoK)
- This is a structured framework developed to systematically evaluate LLM performance in binary-to-source recovery. It involves defining clear research questions and a set of evaluation dimensions derived from the Law of Leaky Abstraction to ensure a holistic assessment.
- Evaluation Dimensions
- These are the specific criteria used to measure how well recovered source code matches the original. They include six metrics covering compilation quality, readability, lexical similarity, syntactic similarity, semantic similarity, and functional recovery.
- Byte-wise Match (BW)
- This is the strictest metric in the study. It requires an exact match of every instruction selected during decompilation. If any divergence occurs between two programs that should be functionally identical, this metric fails.
- Agent4Decompile
- This method, which uses iterative compiler iterations and multi-role LLM feedback, was found to be the best approach for recovering compilable code with byte-level semantics and high functional similarity across various test cases.
Terminology used across episodes
This episode discusses
- SoK: Are LLMs Reliable at Source Code Recovery? A Taxonomy and Empirical Evaluation · Paper Radio
- Evaluating Neural Decompilation of Dart AOT Binaries: Fine-Tuning, Metric Validity, Specification Leakage, and Reliability · Paper Radio
- The Evolution of Binary Decompilation in the Modern Era: A Taxonomy, Literature Review, and Future Perspectives · Paper Radio
- PCodeTrans: Translate Decompiled Pseudocode to Compilable and Executable Equivalent
- Practical Source Code Recovery from Binary Functions Using Anchor-Based Retrieval and LLM Reasoning
- Beyond the C: Retargetable Decompilation using Neural Machine Translation
- Potential and Challenges of Large Language Models for Reverse Engineering · Paper Radio
- CHISEL-ing Back Source Code with AI-enabled Iterative Recovery · Paper Radio
- When LLM Decompilers Recompile More and Preserve Less · Paper Radio
- Binary Decompilation LLM with Feedback-Driven Multi-Turn Refinement
- CodeBLEU: a Method for Automatic Evaluation of Code Synthesis
- Decaf: Improving Neural Decompilation with Automatic Feedback and Search
- Gemma 4 Technical Report
- Statistical Analysis of Executability and Program Equivalence in Decompilation for IoT Vulnerability Detection · Paper Radio
- Constraint-Guided Multi-Agent Decompilation for Executable Binary Recovery
- FidelityGPT: Correcting Decompilation Distortions with Retrieval Augmented Generation
The paper
SoK: Are LLMs Reliable at Source Code Recovery? A Taxonomy and Empirical Evaluation · Read on arXiv
Varun Kohli, Lee Bing Cheng, Nur Hazim Ghazali, Gao Yuze, Daryl Poon, Dinil Mon Divakaran
A*STAR Institute of Advanced Intelligence and Computing, Singapore
Effective source recovery is critical to security applications such as malware analysis, vulnerability assessment, and legacy maintenance. Large Language Models (LLMs) are reshaping this field, shifting the paradigm away from rule-based heuristics to probabilistic and high fidelity semantic recovery of source code from assembly or classical decompiler-derived pseudo-C. However, despite rapid progress, the field suffers from fragmentation across numerous approaches as well as their non-unified evaluations, limiting objective comparisons. Further, existing works have limited coverage of embedded, IoT architectures and source languages beyond C/C++. In this work, we present the first Systematization of Knowledge (SoK) focused specifically on LLM-assisted binary-to-source recovery. We provide a granular design-centric taxonomy of LLM-assisted source recovery methods and systematic evaluations using seven key metrics along six evaluation dimensions. We evaluate state-of-the-art methods on 45,000 test samples derived from four standard and five embedded architectures, five optimization levels, and symbol stripping. We ablate the impact of design choices on recovery performance, including input representation, contextual enrichment, model scale, iteration and review roles using controlled in-house recovery pipelines and three off-the-shelf models. Finally, we test same-language and cross-language recovery capability covering four mature and two legacy languages. Our systematization and comprehensive evaluations provide key insights that guide future directions in this field.
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "SoK: Are LLMs Reliable at Source Code Recovery? A Taxonomy and Empirical Evaluation".
Elias: The gist The work presents the first Systematization of Knowledge (SoK) focused specifically on LLM-assisted binary-to-source recovery,
Nadia: First, who's behind it and why it matters.
Paper summary: Nadia: So looking at this whole thing, the authors of "SoK: Are LLMs Reliable at Source Code Recovery? A Taxonomy and Empirical Evaluation" are essentially giving us a blueprint for how to properly assess these LLM-assisted recovery tools. It’s not about finding one magic solution but creating a structured way to map out all the approaches.
Elias: Right, and the name of this work itself, SoK, points directly to that idea—Systematization of Knowledge—which is what they claim they are doing by defining these four research questions and six evaluation dimensions >
Priya: What this means for us in practice is that when someone looks at a new tool claiming to recover source code, they shouldn't just trust the flashy performance numbers; they need to ask if it’s been tested against the same kinds of scenarios—like different architectures or optimization levels—that this paper laid out >
Nadia: The implication is that we need more standardized benchmarks so we can actually compare these competing techniques objectively, instead of just seeing which one produces the prettiest output >
Elias: They show that while Agent4Decompile is leading in certain areas, the performance drops when you move to specific hardware constraints or aggressive compiler settings, which tells us where those tools are currently brittle >
Priya: It also shows that cross-language recovery to C is better than trying to get back the original source language directly because the AI has less context when dealing with those different runtimes >
Nadia: So ultimately, this paper gives us a framework for future research, showing us exactly which design choices—like using iterative refinement or specific input representations—tend to lead to better results in terms of both correctness and human readability >
Conclusion: Nadia: So we’ve been looking at this new work, "SoK: Are LLMs Reliable at Source Code Recovery? A Taxonomy and Empirical Evaluation," and what we really need to understand is how these AI tools actually stack up against real code recovery tasks.
Elias: Yeah, the authors are building a whole system here—a taxonomy—to sort through all the different ways we try to pull source code back from binary files using LLMs. It's trying to give us a clear map instead of just throwing random tools at the problem.
Priya: From what I see in their setup, they’re not just testing one thing; they’re setting up four main research questions to evaluate different recovery methods across various architectures and optimization levels. That sounds like a pretty thorough way to test something so complex.
Nadia: It seems like the core idea is defining exactly what we mean when we say a recovery method is good or bad by creating these specific evaluation dimensions for quality, similarity, and functional results.
Elias: Exactly, they’re using this six-dimension suite—things like byte-level matches and functional re-executability—to give us a very strict way to measure the fidelity of the recovery. It makes you really look at the actual output rather than just how "smart" the model sounds.
Priya: And what’s interesting is how they test it across different languages, including some legacy ones and cross-language scenarios, which tells us where these AI models are strongest and weakest when dealing with different code structures.
Nadia: They found that one particular method, Agent4Decompile, actually does quite well when you're looking for byte-level accuracy and functional similarity in the recovered code. But there are still clear performance drops on specific hardware like Thumb2 or when the compiler is really aggressive with optimizations.
Elias: That’s a key caveat they point out; performance dips under certain conditions, which tells us those models aren't always robust across every single scenario we might encounter in real-world recovery.
Priya: So for someone just listening to the show, what this means is that these AI recovery tools aren't all equal; you need to know *how* they were trained and *what* constraints they’re operating under to judge if the resulting source code is actually trustworthy.
Nadia: It really boils down to moving away from just checking if the AI can guess what the code *should* look like, toward a systematic way of measuring how accurate that guess is at the instruction level.
Elias: And this whole paper sets up a framework for what future researchers should be looking at when they try to build better tools for this kind of problem. It’s a blueprint, not just another result.
Priya: So they’ve given us the map and the measuring stick, which means now we can start asking much smarter questions about making these kinds of recovery systems more reliable and trustworthy in the future.
More episodes
- 2610.10597-Certified Corruption Budgets: Anytime-Valid Leaderboard Claims under Adaptive Rigging
- 2610.10608-From Investigation Failures to Reliable SOC Agents: Understanding and Improving LLM-Based Alert Triage
- 2610.10612-PyCache Trap: The Inspection-Execution Gap in Agent Skill Scanners
- 2610.10644-SoK: Failure Modes in Common Criteria Product Evaluation - A Taxonomy and Design-for-Evaluability Guidance
- 2610.10617-MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking
- 2610.10620-When AI Finds Hidden Messages, Does It Report?
- 2610.10625-Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
- 2610.10992-The Hint Weight of ML-DSA Signatures Is Key-Dependent: An Empirical Study across the Three FIPS 204 Parameter Sets
- 2610.10659-Applying Security by Design at the Point of Execution: How Governed Security Requirements Affect the Security of AI-Generated Code
- 2610.10735-DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits