SoK: Are LLMs Reliable at Source Code Recovery? A Taxonomy and Empirical Evaluation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "SoK: Are LLMs Reliable at Source Code Recovery? A Taxonomy and Empirical Evaluation".
Elias: The gist The work presents the first Systematization of Knowledge (SoK) focused specifically on LLM-assisted binary-to-source recovery,
Nadia: First, who's behind it and why it matters.
Paper summary: Nadia: So looking at this whole thing, the authors of "SoK: Are LLMs Reliable at Source Code Recovery? A Taxonomy and Empirical Evaluation" are essentially giving us a blueprint for how to properly assess these LLM-assisted recovery tools. It’s not about finding one magic solution but creating a structured way to map out all the approaches.
Elias: Right, and the name of this work itself, SoK, points directly to that idea—Systematization of Knowledge—which is what they claim they are doing by defining these four research questions and six evaluation dimensions >
Priya: What this means for us in practice is that when someone looks at a new tool claiming to recover source code, they shouldn't just trust the flashy performance numbers; they need to ask if it’s been tested against the same kinds of scenarios—like different architectures or optimization levels—that this paper laid out >
Nadia: The implication is that we need more standardized benchmarks so we can actually compare these competing techniques objectively, instead of just seeing which one produces the prettiest output >
Elias: They show that while Agent4Decompile is leading in certain areas, the performance drops when you move to specific hardware constraints or aggressive compiler settings, which tells us where those tools are currently brittle >
Priya: It also shows that cross-language recovery to C is better than trying to get back the original source language directly because the AI has less context when dealing with those different runtimes >
Nadia: So ultimately, this paper gives us a framework for future research, showing us exactly which design choices—like using iterative refinement or specific input representations—tend to lead to better results in terms of both correctness and human readability >
Conclusion: Nadia: So we’ve been looking at this new work, "SoK: Are LLMs Reliable at Source Code Recovery? A Taxonomy and Empirical Evaluation," and what we really need to understand is how these AI tools actually stack up against real code recovery tasks.
Elias: Yeah, the authors are building a whole system here—a taxonomy—to sort through all the different ways we try to pull source code back from binary files using LLMs. It's trying to give us a clear map instead of just throwing random tools at the problem.
Priya: From what I see in their setup, they’re not just testing one thing; they’re setting up four main research questions to evaluate different recovery methods across various architectures and optimization levels. That sounds like a pretty thorough way to test something so complex.
Nadia: It seems like the core idea is defining exactly what we mean when we say a recovery method is good or bad by creating these specific evaluation dimensions for quality, similarity, and functional results.
Elias: Exactly, they’re using this six-dimension suite—things like byte-level matches and functional re-executability—to give us a very strict way to measure the fidelity of the recovery. It makes you really look at the actual output rather than just how "smart" the model sounds.
Priya: And what’s interesting is how they test it across different languages, including some legacy ones and cross-language scenarios, which tells us where these AI models are strongest and weakest when dealing with different code structures.
Nadia: They found that one particular method, Agent4Decompile, actually does quite well when you're looking for byte-level accuracy and functional similarity in the recovered code. But there are still clear performance drops on specific hardware like Thumb2 or when the compiler is really aggressive with optimizations.
Elias: That’s a key caveat they point out; performance dips under certain conditions, which tells us those models aren't always robust across every single scenario we might encounter in real-world recovery.
Priya: So for someone just listening to the show, what this means is that these AI recovery tools aren't all equal; you need to know *how* they were trained and *what* constraints they’re operating under to judge if the resulting source code is actually trustworthy.
Nadia: It really boils down to moving away from just checking if the AI can guess what the code *should* look like, toward a systematic way of measuring how accurate that guess is at the instruction level.
Elias: And this whole paper sets up a framework for what future researchers should be looking at when they try to build better tools for this kind of problem. It’s a blueprint, not just another result.
Priya: So they’ve given us the map and the measuring stick, which means now we can start asking much smarter questions about making these kinds of recovery systems more reliable and trustworthy in the future.
Varun Kohli, Lee Bing Cheng, Nur Hazim Ghazali, Gao Yuze, Daryl Poon, Dinil Mon Divakaran
A*STAR Institute of Advanced Intelligence and Computing, Singapore
cs.CR, cs.SE
Submitted: 2026-10-08
Updated: 2026-10-08
License: http://creativecommons.org/licenses/by/4.0/
The gist: The gist The work presents the first Systematization of Knowledge (SoK) focused specifically on LLM-assisted binary-to-source recovery, providing a granular design-centric taxonomy and systematic
Key concepts
- Systematization of Knowledge (SoK)
- This is a structured framework developed to systematically evaluate LLM performance in binary-to-source recovery. It involves defining clear research questions and a set of evaluation dimensions derived from the Law of Leaky Abstraction to ensure a holistic assessment.
- Evaluation Dimensions
- These are the specific criteria used to measure how well recovered source code matches the original. They include six metrics covering compilation quality, readability, lexical similarity, syntactic similarity, semantic similarity, and functional recovery.
- Byte-wise Match (BW)
- This is the strictest metric in the study. It requires an exact match of every instruction selected during decompilation. If any divergence occurs between two programs that should be functionally identical, this metric fails.
- Agent4Decompile
- This method, which uses iterative compiler iterations and multi-role LLM feedback, was found to be the best approach for recovering compilable code with byte-level semantics and high functional similarity across various test cases.
Terminology
Summary
The gist The work presents the first Systematization of Knowledge (SoK) focused specifically on LLM-assisted binary-to-source recovery, providing a granular design-centric taxonomy and systematic evaluations to guide future directions in this field
Taxonomy and Evaluation Framework
The paper introduces a Systematization of Knowledge (SoK)
dedicated to LLM-assisted source code recovery, formulating four core Research Questions (RQs) to evaluate state-of-the-art (SOTA) decompilation paradigms
These RQs are: RQ1: Evaluation Dimensions, RQ2: Generalization, RQ3: Pipeline Design, and RQ4: Multi-Language Recovery Spectrum
The evaluation dimensions are derived from the Law of Leaky Abstraction, considering what dimensions must be considered to evaluate the fidelity of source recovery
The six key evaluation dimensions selected for a holistic evaluation are two quality metrics and four similarity metrics
Empirical Evaluation Scope
The study provides a comprehensive evaluation of SOTA methods on 45,000 test samples derived from four standard and five embedded (IoT) architectures, five optimization levels, and symbol stripping
The evaluation covers various aspects of design choices through controlled ablations using three off-the-shelf Gemma4 models (12b-it, 31b-it) and Qwen3.8-27b models
The evaluation also tests same-language and crosslanguage recovery capability covering four mature and two legacy languages
Key Findings on Performance
The key findings indicate that Agent4Decompile is the best method for recovering compilabile code with byte-level semantics and functional similarity
SOTA methods perform worse on Thumb2 (IoT) ISAs than standard ISAs, and performance drops on aggressive optimization, while symbol stripping mainly impacts readability
Pseudo-C acts as a portability layer for cross-architecture recovery due to standardized output of classical decompilers
Iterative compiler iterations and larger models improve compilability, binarylevel semantics and functional recovery, while LLM roles improve readability
Metric Suite Definition
The paper defines a holistic evaluation suite based on six dimensions: Compilation Quality (Re-compilation Rate [RC], Compile-error Distribution), Readability Quality (Relative Readability Index [R2I]), Lexical Similarity (CodeBLEU, Edit Similarity), Syntactic Similarity (CodeBLEU AST-match), Semantic Similarity (CodeBLEU DF-match, Byte-wise Match [BW]), and Functional Similarity (Re-executability Rate)
Byte-wise match (BW) is the strictest metric in their suite since any divergence in instruction selection fails it, even between behaviorally identical programs and a pass certifies recovery up to compiler determinism
Impact of Design Choices
Methods that ingest pseudo-C are shown to outperform those that use assembly across all metrics, architectures, optimizations, and symbol variants
Iterative multi-role methods like Agent4Decompile report the highest RCI and BW compared to all other control flows through their targeted refinement iterations using in-loop compiler and functional feedback
The impact of model scale is also significant: larger LLMs achieve higher overall fidelity, and methods that use general-purpose, off-the-shelf LLMs match the performance of those fine-tuned on pseudo-C
Cross-Language Recovery
The fidelity of source recovery is worse on non-C derived binaries, but cross-recovery to C is better than recovering the original source
Single-pass RCI for cross-language recovery is far lower because the model must strip the source language’s runtime scaffolding from the pseudo-C
Iteration repairs lead to nearly 2× functional gains for Fortran (17%→29%) and C++ (14%→27%), while Go and Rust see poor performance since their binaries embed goroutine scheduling, bounds-check panics, and garbage-collection machinery that the LLM is unable to translate into plain C<ref:
Improvements for AI systems
-
Improved Systematization of Knowledge (SoK) for LLM-assisted source recovery: The SoK provides a
granular design-centric taxonomy of LLM-assisted source recovery methods,
allowing for asystematic evaluation of SOTA and targeted ablations
across five orthogonal axes (model setup, control flow, number of model roles per turn, input representations, and augmentation). -
Comprehensive Evaluation Metrics Suite: The paper proposes a
holistic categorization of evaluation metrics for source recovery into the two objectives and six dimensions,
selecting seven key metrics (Re-compilation Rate with caller Interface (RCI)
andRe-execution Rate (RE)
for quality/functional similarity, plus measures for readability, lexical, syntactic, semantic, and functional similarity). -
Taxonomy-Driven Ablation Framework: The methodology allows for a
controlled, taxonomy-driven ablation using three offthe shelf Gemma4 (12b-it, 31b-it) [67] and Qwen3.8-27b [52] models, various base input representations and enrichment channels,
enabling researchers to systematically measure the impact of specific design choices on recovery performance. -
Architecture and Optimization Generalization Testing: The system can perform
Generalization
by evaluating methods againstcross-architecture binaries covering standard and embedded (IoT) architectures, diverse compiler optimization flags (O0-Os), and symbol availability,
specifically demonstrating performance drops on Thumb2 ISAs versus standard ISAs. -
Non-C Language Recovery Capability Assessment: The system can assess
Non-C recovery
by evaluating the capability of current LLM and classical decompiler capabilities to handle languages beyond C/C++, highlighting thatCurrent LLM and classical decompiler capabilities are C-centric and heavily under-perform in non-C recovery.
Abstract
Effective source recovery is critical to security applications such as malware analysis, vulnerability assessment, and legacy maintenance. Large Language Models (LLMs) are reshaping this field, shifting the paradigm away from rule-based heuristics to probabilistic and high fidelity semantic recovery of source code from assembly or classical decompiler-derived pseudo-C. However, despite rapid progress, the field suffers from fragmentation across numerous approaches as well as their non-unified evaluations, limiting objective comparisons. Further, existing works have limited coverage of embedded, IoT architectures and source languages beyond C/C++. In this work, we present the first Systematization of Knowledge (SoK) focused specifically on LLM-assisted binary-to-source recovery. We provide a granular design-centric taxonomy of LLM-assisted source recovery methods and systematic evaluations using seven key metrics along six evaluation dimensions. We evaluate state-of-the-art methods on 45,000 test samples derived from four standard and five embedded architectures, five optimization levels, and symbol stripping. We ablate the impact of design choices on recovery performance, including input representation, contextual enrichment, model scale, iteration and review roles using controlled in-house recovery pipelines and three off-the-shelf models. Finally, we test same-language and cross-language recovery capability covering four mature and two legacy languages. Our systematization and comprehensive evaluations provide key insights that guide future directions in this field.
Sources
- Evaluating Neural Decompilation of Dart AOT Binaries: Fine-Tuning, Metric Validity, Specification Leakage, and Reliability
- The Evolution of Binary Decompilation in the Modern Era: A Taxonomy, Literature Review, and Future Perspectives
- PCodeTrans: Translate Decompiled Pseudocode to Compilable and Executable Equivalent
- Practical Source Code Recovery from Binary Functions Using Anchor-Based Retrieval and LLM Reasoning
- Beyond the C: Retargetable Decompilation using Neural Machine Translation
- Potential and Challenges of Large Language Models for Reverse Engineering
- CHISEL-ing Back Source Code with AI-enabled Iterative Recovery
- When LLM Decompilers Recompile More and Preserve Less
- Binary Decompilation LLM with Feedback-Driven Multi-Turn Refinement
- CodeBLEU: a Method for Automatic Evaluation of Code Synthesis
- Decaf: Improving Neural Decompilation with Automatic Feedback and Search
- Gemma 4 Technical Report
- Statistical Analysis of Executability and Program Equivalence in Decompilation for IoT Vulnerability Detection
- Constraint-Guided Multi-Agent Decompilation for Executable Binary Recovery
- FidelityGPT: Correcting Decompilation Distortions with Retrieval Augmented Generation
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs