Potential and Challenges of Large Language Models for Reverse Engineering

arXiv:2509.21821 · cs.CR · Submitted 2025-09-26 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.

Elias: Today's paper: "Potential and Challenges of Large Language Models for Reverse Engineering".

Nadia: Reverse Engineering (RE) remains a labor-intensive process central to software security,

Elias: First, who's behind it and why it matters.

Paper summary: Nadia: So, to summarize this paper, "Potential and Challenges of Large Language Models for Reverse Engineering," the main thesis is that while LLMs are being applied to reverse engineering for things like vulnerability discovery and malware analysis, their role compared to previous machine learning is still unclear because some efforts are just adapting existing pipelines with minimal changes while others are exploring broader reasoning abilities.

Elias: That points to a fundamental difference in approach; it suggests we're not seeing one single way LLMs are being leveraged for RE, but rather a spectrum of applications driven by different goals and capabilities.

Priya: The paper makes the claim that there's a significant gap between low-level code and high-level reasoning that LLMs are trying to close, and they stress that this gap is causing fragmentation in how these models are being used across the field.

Nadia: Precisely; the paper highlights this tic gap as a major issue, showing that without some consolidation, we're seeing task-specific models evaluated by totally disparate methods and lacking any common benchmarks to compare their actual performance.

Elias: It’s concerning that this lack of a unified framework hinders cumulative knowledge building, which is something I worry about when thinking about long-term cryptographic analysis or deep security research.

Priya: Furthermore, they note that the divergence in assumptions between open-source implementations and academic studies means there's a real gap between what people are imagining conceptually and what can actually be deployed in a practical setting.

Nadia: The paper essentially argues that to move forward effectively, we need to address this fragmentation by providing a systematic mapping of all existing LLM applications in RE. This mapping allows us to organize the landscape by objective, target, method, evaluation strategy, and data scale.

Elias: That systematic approach is what makes the paper matter; it moves us away from just looking at isolated experiments and gives us a way to compare different approaches on a more level playing field.

Priya: It seems like this mapping effort is crucial because it helps illuminate the different ways these models are being used, which is vital for understanding where they can actually be applied safely and effectively.

Conclusion: Nadia: Looking at the title, "Potential and Challenges of Large Language Models for Reverse Engineering," it tells us that this isn't just a celebration of what LLMs can do, but rather a serious look at both the opportunities and the significant hurdles we face when trying to use these tools in security analysis.

Elias: I agree; the paper by Hu et al. is important because it moves beyond simply listing cool applications and instead focuses on structuring the entire research ecosystem around LLMs in RE so we can actually assess their real impact.

Priya: From a measurement standpoint, the implication is that we need better ways to evaluate these models consistently, since they’ve pointed out that different evaluation methods obscure whether we are seeing genuine improvements or just superficial changes in performance.

Nadia: Exactly; if we can use this proposed five-dimensional taxonomy, it should help us decide which kinds of tasks—like performance improvement versus interpretability—are most viable right now for security teams to actually adopt.

Elias: And from a cryptographic viewpoint, if the authors manage to bring some order here, it might help us better understand the assumptions these models make when handling code representations, which is something we need to scrutinize closely.

Priya: Ultimately, the implication for the wider field is that we need consolidated frameworks so that knowledge can accumulate more predictably and responsibly as people start applying this technology to security-critical tasks.

Nadia: So, in simple terms, this paper provides a map of where LLMs are in reverse engineering right now, helping us see what's working and what's causing confusion about the field.

Elias: It’s a necessary step toward making sure that when we start using these powerful generative tools for security analysis, we have a clear understanding of their limits and their actual potential.

School of Information Studies, McGill University

cs.CR

Submitted: 2025-09-26

Updated: 2026-10-03

Code: https://github.com/atredispartners/aidapal

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 71/100

The gist: Reverse Engineering (RE) remains a labor-intensive process central to software security, and this paper systematizes the rapidly evolving field of applying Large Language Models (LLMs) to RE by

Key concepts

Objective
This dimension defines the core purpose of an LLM application in reverse engineering. It categorizes research based on what the study aims to achieve, such as improving performance metrics like analysis speed, enhancing code readability through interpretability, or using LLMs for novel discovery of vulnerabilities.
Target Representations
This classifies what the LLM is analyzing at different levels of abstraction. Targets range from raw binary sequences and assembly code to decompiled source code, showing how models are applied across various stages of program analysis.
Methodological Strategies
This describes the 'how' an LLM interacts with reverse engineering tasks. Methods include simple zero/few-shot prompting for rapid prototyping, fine-tuning for high precision like decompilation, and agent-based systems that integrate LLMs with external tools like debuggers.
Data Scale
This dimension measures the volume of data used to train or test the LLM applications. It ranges from minimal proof-of-concept examples to fine-tuning on moderate datasets, up to massive pre-training on billions of tokens for broad representational capacity.

Terminology

Summary

Reverse Engineering (RE) remains a labor-intensive process central to software security, and this paper systematizes the rapidly evolving field of applying Large Language Models (LLMs) to RE by reviewing 44 research papers and 18 open-source projects. This work is significant because it proposes a five-dimensional taxonomy that organizes existing LLM applications in RE by objective, target, method, evaluation strategy, and data scale to provide a unified framework for comparison.

The gist

This paper provides the first systematic mapping of LLM applications in RE by proposing a five-dimensional taxonomy that categorizes studies by objective, target, method, evaluation strategy, and data scale.

Taxonomy Dimensions

The research is organized around five complementary dimensions to structure the field:

  1. Objective: This dimension captures the fundamental why of the core research question or practical problem. The categories include:

  2. Performance: Focusing on improving efficiency or effectiveness of established RE tasks, emphasizing metrics such as analysis speed, accuracy, and false-positive reduction.

  3. Interpretability: Emphasizing rendering code comprehensible to human analysts by generating meaningful variable and function names or producing natural language summaries.

  4. Discovery: Addressing the pursuit of novel insights, positioning LLMs as exploratory instruments for hypothesis generation for tasks like detecting previously unreported vulnerabilities.

  5. Robustness: Concerned with the reliability of systems in adversarial environments, addressing challenges such as defending against prompt injection attacks and mitigating hallucinations in safety-critical analysis.

Target Representations

The Target dimension addresses the “what” by classifying entities subjected to LLM-based analysis along a continuum of abstraction:

  1. Raw Bytes: Analyzing the raw binary sequences, including tasks like identifying file types based on magic bytes and detecting obfuscation through entropy-based patterns.

  2. Assembly Code: Focusing on recovering program semantics from instruction-level representations, such as inferring function prototypes and recognizing calling conventions.

  3. Decompiled Code: The product of decompilers, LLMs are predominantly applied here for Interpretability tasks like renaming variables and restating logic.

  4. Source Code: Analyzing artifacts like assembly or decompiled code, requiring models to capture both linguistic patterns and formal program logic.

Methodological Strategies

The Method dimension delineates the “how,” reflecting a trade-off between generality and specificity:

  1. Zero/Few-Shot Prompting (Z/FSP): Interacting directly with pre-trained LLMs through prompts, emphasizing rapid prototyping and leveraging general knowledge embedded in the model.

  2. Fine-Tuning: Adapting LLMs to RE-specific data through supervised training or parameter-efficient tuning methods, used for tasks where precision is critical, such as binary-to-source decompilation.

  3. Retrieval-Augmented Generation (RAG): Enhancing LLMs with external knowledge bases like disassembly manuals to improve factual accuracy and diminish hallucination.

  4. Agent-Based: Conceptualizing the LLM as a component in a multi-agent system that integrates external tools like debuggers, symbolic execution engines, and disassemblers, aiming to emulate the workflow of a human analyst.

  5. Data Generation and Processing (Data G&P): Leveraging LLMs to synthesize, augment, or preprocess data for RE tasks to address the chronic scarcity of labeled RE data.

Evaluation and Data Scale

The Evaluation dimension assesses effectiveness using:

  1. Expert-Based Assessment: Relying on the judgment of human evaluators to assess readability of decompiled code or the plausibility of vulnerability explanations.

  2. Automated Metric Scoring: Using quantitative measures like BLEU, ROUGE, and perplexity for text generation, which enables scalability and reproducibility.

  3. Ground-Truth Validation: Evaluating outputs against authoritative reference labels or established benchmarks, providing a rigorous but resource-intensive benchmark.

The Data Scale dimension characterizes the scope of data used:

  1. Proof-of-Concept / Few-Shot: Operating on minimal data, often limited to a small number of prompt examples to demonstrate feasibility.

  2. Fine-Tuning Scale: Involving moderate-sized datasets, typically ranging from thousands to millions of labeled samples, achieving measurable task alignment.

  3. Pre-Training / Massive Scale: Training or adapting models on billions of tokens drawn from diverse sources, which aims to endow models with broad representational capacity.

Analysis and Gaps

The quantitative landscape reveals a strong predominance of performance-oriented studies (96.77%), while interpretability (11.29%), discovery (8.

Improvements for AI systems

Based on the comprehensive systematization presented in this State-of-Knowledge (SoK) paper, here are specific, actionable improvements for existing or future AI systems applied to Reverse Engineering (RE), categorized by how they map onto the proposed taxonomy:


The improved AI system will move beyond simple black-box automation and become a multi-faceted, robust security reasoning engine capable of navigating the entire spectrum from low-level binary manipulation to high-level code synthesis.

Here are the specific improvements:

The improved AI system will be able to perform the following specific tasks:

  1. To achieve superior performance in identifying subtle, obfuscated vulnerabilities (e.g., 0-day patterns) by integrating specialized fine-tuning on high-quality, curated datasets, while simultaneously maintaining reliability against adversarial prompt injection attacks.

  2. To enhance human comprehensibility by producing highly accurate decompiled code with semantically correct variable and function names and detailed annotations, acting as an intelligent translator between assembly instructions and higher-level programming concepts.

  3. To perform novel exploratory analysis by generating hypotheses for undocumented protocols or latent defects that escape standard testing, moving beyond known exploit patterns to discover previously unseen software behaviors.

  4. To operate reliably in complex, real-world environments by dynamically integrating external tools (debuggers/symbolic execution engines) via an agent-based framework, allowing the model to iteratively refine its analysis based on runtime feedback and ensuring outputs are grounded in factual knowledge retrieved from authoritative documentation (RAG).

  5. To provide a transparent and auditable analysis pipeline by outputting not just results, but also comprehensive rationales for every prediction, uncertainty estimates for every claim, and clear links to the specific data/evidence used to ground its reasoning.

Abstract

Reverse engineering (RE) is central to cybersecurity, supporting tasks such as decompilation, deobfuscation, and security analysis. However, RE remains labor-intensive and expertise-demanding, as analysts often manually recover high-level semantics from low-level program representations. Recent advances in large language models (LLMs) provide a promising way to address these challenges through program artifact understanding, semantic reasoning, and tool-augmented problem solving. This potential has stimulated interest in LLM-assisted RE, but existing studies remain scattered across different tasks, targets, methodologies, and evaluation practices. Despite these advances, the literature still lacks a comprehensive survey that consolidates progress, systematizes technical choices, and clarifies open challenges and future opportunities. To fill this gap, we present a systematic survey of LLM-assisted RE, covering 48 peer-reviewed and published research articles identified through our search and selection protocol as of July 1, 2026. We develop a faceted taxonomy that organizes prior studies along six dimensions: task objective, analysis target, methodological approach, evaluation protocol, training scale, and data quality. We extract task formulations and experimental settings from existing studies to support comparison, reproducibility, and future research. From this review, we synthesize research gaps, characterize key challenges, and outline future directions toward more reliable, reproducible, and security-relevant applications of LLMs in RE.

Sources

Related papers