Instruction Alignment for Binary Code Representation Learning

arXiv:2608.11766 · cs.SE, cs.AI, cs.CR · Submitted 2026-08-12 · Read on arXiv

Huaijin Wang, Shuai Wang

Shandong University · Hong Kong University of Science and Technology

cs.SE, cs.AI, cs.CR

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: In proceedings of the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE 2026)

Code: https://github.com/whj0401/InsnAlign

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 75/100

The gist: This paper proposes InsnAlign, a training approach that leverages instruction-level alignment knowledge to improve binary code representation learning.

Terminology

Summary

This paper proposes InsnAlign, a training approach that leverages instruction-level alignment knowledge to improve binary code representation learning. The work is motivated by the observation that existing binary code representation learning methods primarily rely on coarse-grained function-level supervision, such as symbol-based matching between binary functions compiled from the same source under different configurations, while overlooking fine-grained instruction-level correspondences available from compiler debug information.

The authors first conduct a preliminary measurement study using jTrans, a Transformer-based binary code embedding model, to evaluate how well existing models capture instruction-level semantics. They formulate instruction alignment as a retrieval task: given a pair of binary functions compiled from the same source, and given an instruction in one function, they measure the model's ability to retrieve the semantically corresponding instruction in the other function. Their measurement reveals that "models finetuned with function-level contrastive learning exhibit substantially better instruction alignment than their pre-trained version, suggesting a strong correlation between instruction-level understanding and embedding quality." Specifically, the finetuned model consistently outperforms the pre-trained model in both Mean Reciprocal Rank (MRR) and Recall@1 across different optimization level pairs.

Based on this observation, the authors design InsnAlign, which explicitly incorporates instruction alignment as an auxiliary objective alongside function-level contrastive learning. The framework extends existing Transformer-based binary code embedding models, including jTrans and CLAP, by augmenting standard function-level contrastive training with an instruction-level alignment objective derived from fine-grained source-line correspondences. Given a pair of binary functions compiled from the same source, the framework constructs an instruction alignment matrix using debug information, encodes both functions with a shared Transformer, derives instruction embeddings by mean-pooling token-level hidden states, and computes an instruction similarity matrix using cosine similarity. An InfoNCE loss is then formulated over the alignment matrix and similarity matrix to optimize instruction alignment, jointly combined with the function-level triplet loss.

The instruction alignment loss uses a multi-positive generalization inspired by supervised contrastive learning, since instruction alignment is inherently many-to-many—a single source line may expand into multiple instructions, and multiple lines may share instructions. The loss is computed symmetrically in both directions (A→B and B→A). The combined training objective is L = Ltriplet + λ·Lalign, where λ is a weighting coefficient. The authors adopt a layer freezing strategy, freezing the embedding layer and the first L encoder layers of the BERT backbone during training to reduce computational cost while preserving pre-trained knowledge.

For training data, the authors reuse the BinaryCorp dataset's training partition, using only the subset containing debug information, which contains 1,655,011 binary functions from 1,544 projects, forming 2,326,328 positive function pairs. For evaluation, they construct a test set from BinKit, manually confirming that none of the BinKit projects overlap with BinaryCorp projects to ensure clean evaluation without data leakage. They follow a deduplication preprocess from vSim and remove 18 projects that exist in the BinaryCorp training set.

The evaluation addresses five research questions:

RQ1 (Training Loss Convergence): The instruction alignment loss drops sharply during the first epoch when λ > 0, while the function-level contrastive loss remains near zero throughout training, confirming that the auxiliary loss does not degrade existing function-level similarity capability. Models trained with the auxiliary task significantly outperform those without by 50.9% for jTrans and 88.2% for CLAP on Recall@1 for instruction alignment.

RQ2 (Impact on Function Embedding): Instruction alignment training improves function-level Recall@1 across all 16 cross-compiler, cross-optimization settings (4 compilers at O0 vs. O3). For jTrans, the average Recall@1 improves from 0.4011 to 0.4282; for CLAP, from 0.6396 to 0.6590. The improvement holds consistently across both recent and outdated compilers, demonstrating robustness to distribution shift.

RQ3 (Discriminability of Similar Functions): The authors define a Mean Alignment Score (MAS) that aggregates instruction-level similarities into a single scalar. Under hard negative sampling (top-5 candidates by function-level cosine similarity), MAS attains a higher Cohen's d than cosine similarity for every model and a higher AUC in most cases, indicating that the instruction-level signal is more discriminative when function embeddings are close. Instruction alignment training improves both AUC and Cohen's d for both MAS and cosine similarity. A synergy score combining function-level cosine similarity and MAS further improves retrieval: for InsnAlignjtrans, synergy scoring increases Recall@1 from 0.4282 to 0.5475 (27.9% improvement), while InsnAlignclap gains 2.17%.

RQ4 (Label Quality and Noise Resilience): The authors assess compiler-generated debug labels for coverage and correctness. Across training data, 86.6% of instructions in-O0 functions are covered in their-O3 counterparts. Manual and LLM assessments classify mappings as CORRECT (86.8%), PLAUSIBLE (5.3%), UNVERIFIABLE (5.1%), or SUSPICIOUS or WRONG (2.8%). With only 2.8% classified as SUSPICIOUS or WRONG, the labels appear highly reliable. Noise resilience experiments show that even with 20% injected noise, InsnAlignjtrans achieves 0.634 Recall@1, still 26.0% higher than the baseline.

RQ5 (Effectiveness with Hard Negatives): When the function-level contrastive objective is strengthened with hard negative mining, instruction alignment remains complementary. InsnAlign continues to improve function-level retrieval over the hard-negative baseline, and synergy re-ranking provides further gains. A paired t-test confirms the synergy gain is statistically significant (t = 27.0, p = 5.2 × 10−161).

The case study presents concrete examples illustrating how instruction alignment provides interpretable evidence for BCSA, including resilience to instruction splitting under optimization, robustness to function inlining, and potential for patch presence detection. The authors analyze 216 BCSA failures and group them into four causes: semantics-deprived wrapper queries (162/216), context-window truncation (14/216), near-twin siblings (20/216), and size/constant/global-only differences (20/216).

The paper's contributions are summarized as: (1) identifying that existing methods only exploit coarse-grained function-level knowledge and overlook instruction-level semantic correspondences; (2) being the first to propose an instruction alignment measurement to evaluate how well binary code embedding models capture fine-grained instruction semantics; and (3) designing a training approach that leverages instruction alignment knowledge to improve binary code embeddings, with extensive experiments demonstrating improved retrieval performance and more discriminative, inspectable evidence for similarity judgments.

Improvements for AI systems

Improvements to AI Systems:

  1. Fine-Grained Semantic Alignment Training for Code Embedding Models
  • Incorporate an auxiliary instruction-level alignment loss (InfoNCE-based, multi-positive) alongside function-level contrastive learning.

  • Use compiler debug information to construct many-to-many instruction correspondence matrices between binary functions.

  • The improved AI system learns embeddings that capture both coarse function semantics and fine-grained instruction semantics, yielding higher retrieval accuracy (e.g., +27.9% Recall@1 with synergy re-ranking) and robustness across compilers and optimization levels.

  1. Instruction-Aware Re-Ranking for Binary Code Similarity
  • After initial function-level retrieval, compute a Mean Alignment Score (MAS) by aggregating instruction-level cosine similarities between candidate and query functions.

  • Combine MAS with function-level cosine similarity via a synergy score for re-ranking.

  • The improved system can disambiguate near-identical functions (hard negatives) with higher discriminative power (higher Cohen’s d and AUC), reducing false positives in tasks like vulnerability search and patch presence detection.

  1. Noise-Resilient Supervision from Compiler Debug Labels
  • Use compiler-generated debug information as training labels, with automated quality filtering (e.g., only keep mappings classified as CORRECT or PLAUSIBLE).

  • Inject controlled noise during training to simulate label imperfections.

  • The improved system maintains high performance even with 20% label noise (Recall@1 = 0.634, still 26% above baseline), making it practical for real-world datasets with imperfect annotations.

  1. Interpretable Similarity Evidence for Binary Code Analysis
  • Output instruction-level alignment maps alongside function-level similarity scores.

  • The improved system provides human-inspectable evidence for why two binaries are similar, aiding analysts in verifying matches, detecting inlined functions, and identifying semantically equivalent but syntactically different code.

  1. Layer-Freezing Strategy for Efficient Training
  • Freeze embedding and lower encoder layers during instruction-alignment fine-tuning.

  • The improved system reduces computational cost while preserving pre-trained knowledge, enabling scalable training on large binary corpora (e.g., 1.65M functions) without sacrificing accuracy.

  1. Cross-Domain Generalization via Instruction-Level Supervision
  • Train on debug-info-rich binaries (e.g., BinaryCorp) and evaluate on unseen projects (e.g., BinKit) with no overlap.

  • The improved system generalizes across compiler versions (recent and outdated) and optimization levels (O0–O3), demonstrating robustness to distribution shift—critical for real-world deployment on diverse firmware and software.

Abstract

Binary code representation learning is a fundamental problem in software security and reverse engineering. Existing methods mainly learn function-level embeddings that capture coarse-grained semantic relationships between binary functions, but they largely ignore fine-grained instruction-level correspondences. This limitation misses valuable supervision signals available from compiler debug information, which can support the learning of more accurate and interpretable binary code representations. We propose to leverage instruction alignment knowledge to further improve binary code representation learning. Our preliminary study reveals that models finetuned for function-level binary code similarity exhibit substantially better instruction alignment than their pre-trained model, suggesting a strong correlation between instruction alignment and function-level embedding quality. Motivated by this observation, we design a training approach that explicitly incorporates instruction alignment as an auxiliary training objective. Our experiments show that instruction alignment training improves retrieval accuracy and provides more discriminative signal for the model's similarity judgments.

Sources

Related papers