LLM-assisted Generation of Pseudo-C2 Servers for IoT Malware Dynamic Analysis

arXiv:2606.21349 · cs.CR · Submitted 2026-06-23 · Read on arXiv

cs.CR

Submitted: 2026-06-23

Updated: 2026-08-20

Journal ref: Proceedings of the 21st Asia Joint Conference on Information Security (AsiaJCIS 2026), EPiC Series in Computing, vol. 111, EasyChair, 2026, pp. 1-16

DOI: 10.29007/7hdm

Code: https://github.com/jgamblin/Mirai-Source-Code

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 82/100

The gist: This paper proposed a pseudo-C2 server automatic generation system that integrates Ghidra with a Large Language Model (LLM), addressing the challenges of short-lived C2 servers and increasing

Terminology

Summary

This paper proposed a pseudo-C2 server automatic generation system that integrates Ghidra with a Large Language Model (LLM), addressing the challenges of short-lived C2 servers and increasing analysis cost in IoT malware dynamic analysis.

The experiments conducted on Mirai demonstrated three key contributions:

  1. The LLM extracted all 20 items of the core protocol elements with 100% accuracy by semantically interpreting control structures such as switch-case statements.

  2. The generated pseudo-C2 server fully reproduced 7 of 10 DDoS attacks (TCP/UDP/GRE) with attack behavior consistent with the original C2.

  3. The resulting analysis platform automates a process that traditionally required advanced reverse engineering skills.

Furthermore, applying the method to a customized variant of the publicly available Mirai source code (Experiment II) succeeded end-to-end, demonstrating that the LLM infers specifications from binary structures without relying on pre-trained data. The significance of this study is highlighted as extending LLM applicability from static analysis assistance to automatic generation of attack infrastructure for dynamic analysis, suggesting applicability to novel IoT malware with previously unseen specifications.

However, the paper also details several limitations and areas for future work:

Limitations and Scope:

  • Observation Failure: The initial experiments were conducted in a closed network that did not contain a responsive DNS server or web server, resulting in observation failures for DNS/HTTP/STOMP attacks. It is noted that the coincidence of these failure attacks between Experiments I and II supports the interpretation that the observation failure stems from the environment side rather than from the method side.

  • Applicability Scope: While effective for Mirai, the applicability to other IoT malware families (Gafgyt/BASHLITE, etc.) and to samples with more sophisticated encrypted communication or custom protocols remains unverified.

  • Quantitative Evaluation: The quantitative evaluation is deemed insufficient because the experiment was confined to short-term verification centered on 10-second attack observations, and verification of long-term connection maintenance (such as keep-alive processing) over hours or days has not been conducted. Additionally, a quantitative comparison with conventional manual analysis regarding the reduction of analysis time remains a challenge.

Future Work:

The authors outline several necessary steps for future research:

  • (i) Verification on other malware families such as Gafgyt (BASHLITE) to enhance generality.

  • (ii) Integrating a virtual honeypot (DNS/web servers) to observe all attack vectors including DNS/HTTP Flood.

  • (iii) Verifying long-term stability and quantitatively evaluating analysis-time reduction.

  • (iv) Fine-tuning open-source LLMs (such as Llama) for malware analysis to ensure confidentiality and reduce dependence on commercial LLMs.

Improvements for AI systems

As a fastidious AI researcher, I have meticulously analyzed this paper, recognizing its significant leap from simple analysis assistance to automated infrastructure generation. The core innovation—the semantic interpretation of binary control structures by an LLM—is powerful.

However, relying on a single two-step process (Ghidra to LLM) is a limitation when facing novel threats or complex proprietary binaries. To improve the current AI system and elevate it from a specialized tool to a generalized platform, I propose the following enhancements:


The original system is robust but brittle. The improvements focus on enhancing input context, generalizing scope, automating verification, and refining LLM reasoning.

Current Limitation: The system relies on GhidraMCP extensions to provide raw byte sequences and basic structure (like the read bytes tool). This is insufficient for complex obfuscation or intricate data flow logic.

Improvement: Integrate a dedicated Data Flow Analysis Module (DFAM) that feeds the LLM not just static definitions, but also dynamic data flow graphs (DFG) derived from Ghidra’s internal state.

  • Mechanism: Before the LLM prompt is generated, a dedicated component traces how specific memory locations (e.g., an encrypted C2 address) are used by various functions to determine if they are consumed or output.

  • Specific Enhancement: Add context regarding register usage and stack manipulation. The LLM needs to know not just what the bytes are, but how they flow through the CPU registers (e.g., this XOR key is loaded into RDI and used as a loop counter).

  • Why this matters: This allows the the LLM to perform semantic reasoning on complex cryptographic routines, rather than relying solely on pattern matching or brute-force decryption, enabling it to handle novel, non-standard encryption schemes.

Current Limitation: The system is designed specifically for IoT botnets (Mirai). Its success is tied to the specific attack vectors and command structures used by Mirai.

Current Limitation: The system's success is measured by comparing the generated C2's output against a known ground truth from the original C2, and manually observing attack behavior in a closed lab environment.

Current Limitation: The LLM is guided by the instruction to extract specifications but has no inherent mechanism for resolving conflicting or ambiguous data found in the binary (e.g., multiple potential C2 endpoints or redundant protocol definitions).

The original system could generate a single, functional C2 for known threats (Mirai). The improved system will be a Generalized Automated Infrastructure Engineer.

  1. Automated Protocol Emulation: It can take any binary file—malware, proprietary industrial software, or network appliances—and generate a fully functional, highly robust emulation layer (a pseudo-C2) that mimics the exact behavior of the target protocol without needing human intervention.

  2. Cross-Platform Deployment: It will not be limited to Python. The PAL/LLM combination allows it to generate C2 implementations in multiple languages and architectures (e.g, a Rust C2 for security, or a simple Go C2 for speed) based on the same abstract specification.

  3. Self-Correcting Analysis: It can analyze novel threats where the communication protocol is intentionally obfuscated or inconsistent, resolving ambiguities logically rather than failing or guessing based on pre-trained data.

  4. Guaranteed Reliability: The APCE ensures that the generated C2 is not just correct, but operationally sound under real-world network conditions, minimizing maintenance and ensuring reliable simulation for dynamic analysis.

Sources

Related papers