Entropy-based Code Adversarial Translation for Real-world Repository Migration

arXiv:2608.09273 · cs.AI, cs.SE · Submitted 2026-08-11 · Read on arXiv

Yushun Tang, Yisen Cao, Zhicheng Chen, Lin Peng, Junkang Mao, Fengyi Song, Yantao Jia

Huawei Technologies Co., Ltd. · The Chinese University of Hong Kong

cs.AI, cs.SE

Submitted: 2026-08-11

Updated: 2026-08-12

Code: https://github.com/yushuntang/ECAT

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

The gist: This paper introduces Entropy-based Code Adversarial Translation (ECAT), a multi-agent framework for automated Android-to-HarmonyOS repository migration.

Terminology

Summary

This paper introduces Entropy-based Code Adversarial Translation (ECAT), a multi-agent framework for automated Android-to-HarmonyOS repository migration. The work addresses the challenge that while LLMs excel at code generation and automated program repair, migrating an entire repository rarely produces a runnable application because long-horizon translation challenges LLM-based agents' ability to maintain repository-level migration objectives.

The paper is authored by Yushun Tang, Yisen Cao, Zhicheng Chen, Lin Peng, Junkang Mao, Fengyi Song, and Yantao Jia from Huawei Technologies Co., Ltd. and The Chinese University of Hong Kong, and was published on arXiv (arXiv:2608.09273v1).

The paper identifies several key challenges in repository migration:

  1. Long-horizon reasoning difficulty: Repository migration requires long-horizon reasoning over large software repositories, where generation errors accumulate and propagate across long-range dependencies.

  2. Scale challenges: As repository size scales from tens to hundreds of thousands of lines of code (LOC), cross-file dependencies and long-range interactions rapidly increase, making globally consistent repository optimization increasingly difficult.

  3. Self-evaluation bias: Existing systems rely on iterative self-refinement where the same agent repeatedly generates and evaluates repository updates, but such self-evaluation often overestimates intermediate repository quality, leading to premature convergence despite unresolved compilation errors, broken dependencies, or missing functionality.

  4. Lack of knowledge accumulation: Migration experience from previous repositories is rarely accumulated, causing similar optimization trajectories to be repeatedly rediscovered from scratch.

ECAT formulates repository migration as an adversarial entropy minimization problem. Given a source Android repository RS = fiS Ni=1, the goal is to migrate it into a target HarmonyOS repository RT. The paper notes that LLM-based repository generation is inherently probabilistic, causing migration errors to accumulate throughout the repository and progressively increase repository-level disorder.

The optimization objective is formalized as:

π* = arg min E[H(RT)]

where π* denotes the optimal migration policy and H(RT) denotes the Code Entropy of the migrated repository.

The paper introduces Code Entropy as a unified repository-level optimization objective score that quantifies the overall disorder of a migrated repository. The authors clarify that We use entropy as an analogy to characterize repository-level uncertainty caused by unresolved migration inconsistencies, rather than a strict information-theoretic entropy.

Code Entropy aggregates migration errors across K complementary dimensions:

  • Static entropy evaluates repository correctness without execution, including compilation, structure, and migration fidelity

  • Dynamic entropy captures runtime defects after deployment on the HarmonyOS emulator, such as crashes and UI inconsistencies

The overall Code Entropy is computed as:

H(Rt) = W⊤E / 1⊤W

where E is the repository entropy vector with normalized scores ei ∈ [0,1] for each dimension, and W contains non-negative importance weights (assigned in the range [0.1, 0.3]).

The paper specifies 14 entropy dimensions (K=14), including: compile, static lint, style, placeholder, app identity, data layer parity, permission parity, page audit, feature check, hallucination, entry nav, source parity, ui align, and runtime liveness. Each dimension is estimated using either deterministic rule-based verification for objectively observable properties or agentic LLM judges for semantic properties.

ECAT employs two specialized agents with distinct responsibilities:

  1. Generator G(·): Performs Android-to-HarmonyOS code generation and repository repair

  2. Discriminator D(·): acts as an adversarial evaluator that challenges the current repository state by exposing high-entropy regions and generating optimization signals

The interaction follows a challenge-and-repair loop:

  • At iteration t, the discriminator evaluates the current repository and produces (Ht, gt) = D(Mt; RS, Rt), where Ht is the estimated Code Entropy and gt is a structured text gradient

  • The generator updates the repository according to Rt+1 = G(gt, Mt; RS, Rt)

  • A candidate repository is accepted only if its Code Entropy decreases: Rt+1 = R̂t+1 if H(R̂t+1) < H(Rt), otherwise Rt+1 = Rt

The paper emphasizes that Both agents are re-instantiated with isolated contexts at every iteration, so neither can access the other's reasoning traces or inherit bias from its own previous judgments.

Since repository optimization is performed in the discrete code space, numerical gradients cannot be propagated through repository states, the discriminator produces structured text gradients that explicitly identify high-entropy regions, diagnose migration failures, prioritize repair targets, and recommend the corresponding skills. An example text gradient includes a summary, a work list with specific files, issues, and suggested skills.

ECAT includes a self-evolving memory tree that serves as the only learnable component continuously updated during deployment. The memory:

  • accumulates reusable migration experience from entropy-reducing optimization trajectories

  • Is organized hierarchically with a root node storing global migration knowledge, intermediate nodes organizing reusable patterns by coarse defect categories, and leaf nodes maintaining generalized migration patterns

  • Discards repository-specific implementation details to maximize transferability across repositories

The memory enables coarse-to-fine retrieval that focuses only on defect-relevant branches and is shared by both the generator and discriminator.

To avoid premature convergence, ECAT uses a sliding-window stopping criterion where Optimization terminates when the repository Code Entropy remains below a predefined threshold ϵ during the most recent l iterations (with l=2 and ϵ=0.01 in experiments).

The paper introduces A2H-RepoBench, the first real-world multiscale repository-level benchmark for Android-to-HarmonyOS migration, containing three representative open-source Android repositories:

  1. Gallery (50K LOC): A multimedia application evaluating user interface migration and media management

  2. AntennaPod (120K LOC): A podcast management application involving multimedia playback and asynchronous task scheduling, presenting substantially more complex cross-module dependencies

  3. Meshtastic (300K LOC): A large-scale mesh communication application supporting Bluetooth communication and hardware interaction, posing significant challenges for long-range dependency preservation

Repository quality is evaluated from two perspectives:

  1. Node Alignment: Measures structural preservation by constructing platform-agnostic semantic graphs using CodeGraph and computing Align = M/Vs, where Vs denotes semantic nodes of the Android repository and M denotes matched node pairs in HarmonyOS.

  2. Agent-as-Judge: Evaluates functional preservation using a predefined feature checklist where each feature is assigned a score of 1, 0.5, or 0 (Full, Partial, or Missing), yielding Agent = (1/N)Σsi.

ECAT consistently achieves the best performance under both metrics across all repositories, improving the average score from 46.4% achieved by the strongest baseline to 74.7%. Key findings:

  • OpenHands and RepoTransAgent suffer from self-evaluation bias

  • ReCodeAgent, though validated independently, lacks an iterative repair loop and attains relatively high Node Alignment but much lower Agent-as-Judge scores because many translated classes preserve the repository structure while containing only shallow implementations or placeholder logic

  • Removing Dynamic Entropy drops the average from 74.7% to 72.7%, mainly on runtime-heavy repositories

Removing the independent discriminator collapses the average score from 74.7% to 28.4%, a larger degradation than removing any other component. The failure amplifies with repository scale:

  • On Gallery: both metrics drop by roughly 40 points as a third of the features are left missing

  • On AntennaPod: 106 of 184 features are judged Partial while only 5 remain Full

  • On Meshtastic: only 7 of 129 features fully functional and 105 missing

The paper explains: the self-evaluating agent systematically overestimates the quality of its own outputs and triggers the stopping criterion prematurely.

With the memory tree, ECAT reaches the stopping criterion in only 38 iterations, compared with 66 iterations without memory on the Gallery repository. This reduces cumulative token consumption from 125M to 65M tokens, demonstrating that transferable migration knowledge accelerates defect localization and repair.

Migration cost scales with repository size: Gallery requires 38 iterations, 12 hours, and 65M tokens; AntennaPod requires 95 iterations, 29 hours, and 180M tokens; Meshtastic requires 126 iterations, 44 hours, and 230M tokens.

ECAT is robust to the choice of base model: all three LLMs reach comparable final quality (83–89% Agent), indicating that the adversarial entropy-minimization loop, rather than the raw capability of a single model, drives migration quality.

The paper summarizes its major contributions as:

  1. Formulation: "We formulate Android-to-HarmonyOS repository migration as an adversarial entropy minimization problem and introduce Code Entropy, a unified repository-level objective that quantifies repository disorder in terms of migration correctness and fidelity."

  2. Framework: We propose ECAT, a generator–discriminator framework that explicitly decouples repository generation from repository evaluation, with successful trajectories distilled into a self-evolving memory tree.

  3. Benchmark: We construct A2H-RepoBench, the first repository-level benchmark for Android-to-HarmonyOS migration, spanning 50K, 120K, and 300K LOC.

The paper acknowledges: "The iterative adversarial optimization inevitably increases inference latency and token consumption compared with one-shot migration. Future work will investigate more efficient optimization strategies and the extension of ECAT to broader repository-level software engineering tasks."

The paper concludes that "By formulating repository migration as an entropy minimization problem, ECAT iteratively reduces repository disorder through adversarial interaction between an independent generator and discriminator, while continuously improving via a self-evolving memory tree that accumulates transferable migration knowledge. Extensive experiments demonstrate that ECAT consistently outperforms existing LLM-based agentic methods on large-scale real-world repositories."

Improvements for AI systems

Improvement 1: Decoupled Generator–Discriminator Architecture with Context Isolation

The improved AI system separates code generation from evaluation using two independent agents with isolated contexts at each iteration. This prevents self-evaluation bias, where a single agent overestimates its own output quality. The system can now:

  • Detect and correct premature convergence in long-horizon tasks (e.g., repository migration, multi-file refactoring) by forcing external validation.

  • Maintain objective quality assessment across thousands of interdependent files, reducing false success signals that plague single-agent iterative refinement.

Improvement 2: Code Entropy as a Unified, Multi-Dimensional Optimization Objective

The system adopts a scalar Code Entropy score aggregating 14 complementary dimensions (static correctness, runtime behavior, UI alignment, feature parity, etc.) with weighted normalization. This enables:

  • Continuous, fine-grained tracking of repository-level disorder during optimization, rather than binary pass/fail checks.

  • Prioritized repair targeting—the system identifies which entropy dimensions (e.g., permission parity vs. ui align) contribute most to overall disorder and allocates generation effort accordingly.

  • Cross-repository transferability: the same objective can be applied to any codebase migration or large-scale refactoring task by reweighting dimensions.

Improvement 3: Self-Evolving Memory Tree for Reusable Migration Knowledge

The system includes a hierarchical memory structure that distills successful entropy-reducing trajectories into generalized, repository-agnostic patterns. This enables:

  • 42% faster convergence (38 vs. 66 iterations) and 48% lower token consumption (65M vs. 125M) on a 50K LOC repository by reusing prior repair strategies.

  • Coarse-to-fine retrieval that skips irrelevant branches, focusing only on defect categories matching the current repository state.

  • Continuous improvement across deployments—the system gets better at migrating new repositories without retraining, as it accumulates patterns for common failure modes (e.g., placeholder logic, broken navigation, missing permissions).

Improvement 4: Adversarial Text Gradients for Discrete Code Space Optimization

Instead of numerical gradients (impossible in discrete code), the discriminator produces structured text gradients that explicitly list high-entropy regions, diagnose root causes, and recommend specific skills. The improved system can:

  • Perform targeted, explainable repairs—each iteration produces a prioritized work list (file, issue, suggested fix) rather than blind regeneration.

  • Avoid destructive edits by accepting only entropy-decreasing updates, ensuring monotonic quality improvement.

  • Handle long-range dependencies (e.g., 300K LOC) by breaking down global disorder into localized, actionable repair signals.

Improvement 5: Sliding-Window Stopping Criterion with Entropy Threshold

The system avoids premature termination by requiring Code Entropy to remain below a threshold (0.01) for two consecutive iterations. This enables:

  • Reliable termination only when the repository is genuinely stable, preventing the false success problem where unresolved compilation errors or missing features are overlooked.

  • Adaptive iteration counts based on repository complexity (38 iterations for 50K LOC, 126 for 300K LOC), scaling gracefully with task difficulty.

Improved AI System Capabilities

The resulting AI system can autonomously migrate or refactor entire software repositories (tens to hundreds of thousands of lines) with:

  • High fidelity: 74.7% average quality score vs. 46.4% for the strongest baseline, with robust performance across different base LLMs (83–89% functional preservation regardless of model choice).

  • Scalability: Handles cross-file dependencies and long-range interactions that break single-agent systems, as shown by 300K LOC repository migration.

  • Efficiency: Learns from past migrations to reduce token consumption and wall-clock time on subsequent tasks.

  • Generalization: Applicable to any platform-to-platform code translation (e.g., iOS to Android, Python to Rust) by redefining the 14 entropy dimensions and retraining the memory tree on new defect categories.

Sources

Related papers