DCVD: Dual-Channel Cross-Modal Fusion for Joint Vulnerability Detection and Localization

arXiv:2605.11015 · cs.CR, cs.AI · Submitted 2026-05-10 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "DCVD: Dual-Channel Cross-Modal Fusion for Joint Vulnerability Detection and Localization".

Elias: The gist: DCVD proposes a unified framework that performs joint function-level detection and statement-level localization by extracting control-dependency and semantic features through two parallel branches,

Nadia: First, who's behind it and why it matters.

Title and authors: Nadia: We started by looking at the title "DCVD: Dual-Channel Cross-Modal Fusion for Joint Vulnerability Detection and Localization." It tells you immediately that this paper isn't just trying to find a vulnerability; it’s also about where exactly that vulnerability is located.

Elias: Right, and the authors are from institutions like Tsinghua University, Hunan University, Dalian Maritime University, The Chinese University of Hong Kong, Shenzhen University, Northwestern Polytechnical University, Shandong University. It's a pretty broad international collaboration.

Nadia: So what that means for us is that this isn't just one lab pushing an idea; it’s a big group working on integrating these different kinds of information sources into one system.

Priya: From my side, I'm curious how they handled the sheer variety of data types—the graph structures, the textual explanations from the LLM, and all that complexity.

Elias: They address that complexity by setting up two parallel branches early on: one for structural features and one for semantic features.

Nadia: And those branches feed into a shared embedding layer, which is where they start preparing those different data types to be compatible with each other.

Priya: That shared embedding sounds like the crucial bridge that lets the structural and semantic information actually communicate with each other meaningfully before fusion happens.

The paper's summary: Nadia: Now, let's talk about what DCVD actually does in terms of its main mechanism. The core idea is to solve the problem where you can find a vulnerability at the function level, but you don't know which specific lines are causing it.

Elias: Exactly. Existing sequence-based methods capture the meaning of tokens well, but they ignore how that code is actually connected structurally in terms of control flow and syntax hierarchies.

Nadia: While graph-based methods get the structure right, they often miss the deep semantic understanding that an LLM can provide about what the code is supposed to be doing functionally.

Priya: So DCVD seems to aim for a system that gets both perspectives simultaneously, using control dependency features from graphs and natural language explanations from an LLM.

Elias: That’s right. Then they introduce the Cross-Modal Fusion module, which uses contrastive learning to align the structural and semantic representations into a unified space.

Nadia: And after that alignment, they use bidirectional cross-attention to let each modality query and attend to the most relevant parts of the other representation, creating a fused feature.

Priya: So what this means in practice is that instead of just getting a structural map or just reading the code text, you get a fused representation that captures both how it’s built and what it’s supposed to do.

The paper's improvements: Elias: Beyond the fusion module, they introduce the Multi-Granularity Supervisor to handle the supervision mismatch between function-level detection and statement-level localization.

Nadia: That supervisor is a two-pronged system with its own branch for function detection and another branch specifically for statement localization.

Priya: I see how that directly addresses the challenge they mentioned in their introduction regarding satisfying both requirements simultaneously, rather than treating one as secondary.

Elias: The function-level branch uses function-level labels to predict vulnerability using a simple MLP to get a binary prediction, ŷ f, optimized with binary cross-entropy loss.

Nadia: And the statement-level branch refines those token representations through self-attention and another MLP to give you that scalar vulnerability score for each individual line.

Priya: So, the improvement here is that they aren't just doing one thing; they are explicitly supervising both granularities using dedicated loss functions for each part of the system.

Conclusion: Nadia: To wrap up this discussion on DCVD, it’s a unified framework that combines multi-source extraction, cross-modal fusion, and multi-granularity supervised learning.

Elias: It really boils down to deep cross-modal fusion between the structural and semantic representations and explicit supervision at both the function and statement levels.

Priya: And what it does is achieve joint vulnerability detection and localization by training both goals collaboratively through a combined loss function.

Nadia: The results on BigVul show that this design choice leads to state-of-the-art performance on both function-level detection and statement-level localization tasks.

Elias: Specifically, they lead in statement-level classification under both the Two-Phase and One-Phase protocols, showing improvements in MCC and F1 scores across those settings.

Priya: It also achieved the best performance on ranking metrics like Top-one MFR, and MAR compared to baselines <ref:2605.11015#pg1>.

Nadia: So this framework is a strong example of how combining different information sources can lead to more robust security analysis tools when you need both high-level detection and low-level pinpointing.

Elias: It’s a solid design that validates the necessity of deep cross-modal fusion for getting those complex tasks done together effectively.

Priya: Overall, it shows that explicit supervision at different levels is a really powerful way to guide an AI system toward doing exactly what you need in security analysis.

Tsinghua University (1) · Hunan University (2) · Dalian Maritime University (3) · The Chinese University of Hong Kong (4) · Shenzhen University (5) · Northwestern Polytechnical University (6) · Shandong University (7) · BNU-HKBU United International College (8) · Sun Yat-sen University (9) · Peng Cheng Laboratory Guangzhou Intelligence Communications Technology Co., Ltd. (10) · The Fifth Electronic Research Institute of MIIT (12)

cs.CR, cs.AI

Submitted: 2026-05-10

Updated: 2026-10-08

Code: https://github.com/vinsontang1/DCVD

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 90/100

The gist: The gist: DCVD proposes a unified framework that performs joint function-level detection and statement-level localization by extracting control-dependency and semantic features through two parallel

Key concepts

Structural Branch
This branch extracts control dependency features from the code's Abstract Syntax Tree (AST) and Control Flow Graph (CFG). It uses Graph Attention Networks (GATs) to understand how different parts of the code are logically connected, providing a structural view of the program.
Semantic Branch
This branch leverages a Large Language Model (LLM) to generate natural language explanations for the code. It extracts deep semantic features by combining token-level code semantics with global functional understanding derived from these LLM explanations.
Contrastive Alignment
This technique is used in the Cross-Modal Fusion module to align the structural and semantic representations into a unified high-dimensional space. By minimizing a contrastive alignment loss, it ensures that features from both modalities are mapped close together, facilitating better interaction.
Multi-Granularity Supervision
Instead of relying on one type of label, DCVD uses two parallel branches with distinct supervision signals: one for function-level detection and another for statement-level localization. This allows the model to be trained effectively at both high and low levels of code detail.

Terminology

Summary

The gist: DCVD proposes a unified framework that performs joint function-level detection and statement-level localization by extracting control-dependency and semantic features through two parallel branches, integrating them via contrastive alignment coupled with bidirectional cross-attention, and introducing explicit supervision signals at both granularities.>

How it works

DCVD is designed to address the limitations of existing approaches that rely on a single information source by performing joint function-level detection and statement-level localization> The framework extracts control dependency and semantic features through two parallel branches: a structure branch and a semantic branch, which operate in parallel> The structure branch employs Graph Attention Networks (GATs) to extract control dependency features from AST and CFG representations, while the semantic branch leverages an LLM to generate natural language explanations of the code and extracts deep semantic features through a shared embedding layer> The resulting structural representation is obtained by element-wise addition of the GAT outputs from the AST and CFG, denoted as F(s) in equation (3)> The semantic representation F(t) incorporates both token-level code semantics and global functional understanding provided by an LLM explanation, denoted as F(t) in equation (4)>

Cross-Modal Fusion

The Cross-Modal Fusion module is introduced to bridge the cross-modal representation gap between structural and semantic features> This module first aligns the two modalities into a unified high-dimensional space via contrastive learning, where global representations gi from F(s) and ci from the semantic branch are aligned using a contrastive alignment loss Lalign (equation 5)> Following alignment, bidirectional cross-attention is employed to enable deep interaction between the two modalities> This involves deriving queries from one modality and keys/values from the other, allowing each modality to selectively attend to the most informative elements of its counterpart> The fused feature H is then computed by combining these interactions: H = σ Wm (6)> Subsequently, a multi-layer Transformer performs deep contextual modeling on H, yielding the enriched representation K in equation (7)>

Multi-Granularity Supervision

To reconcile the supervision asymmetry between function-level detection and statement-level localization, DCVD designs a Multi-Granularity Supervisor> This supervisor comprises two parallel branches: a function-level detection branch and a statement-level localization branch, each equipped with its own supervision signal> The function-level branch leverages function-level labels to determine vulnerability using an MLP gf to produce a binary prediction yˆf (7), optimized via the binary cross-entropy loss Lf> The statement-level branch refines token representations through self-attention and an MLP gs, projecting each token representation to a scalar vulnerability score sl (9), which is then aggregated within each line to yield the line-level vulnerability probability yˆ(l)s>

Training Objective and Results

The total training objective L is formed by combining the function-level loss Lf, the alignment loss Lalign, and the statement-level loss Ls: L = α (Lf + β · Lalign) + (1 − α) · Ls (10)> This design ensures that both granularities are jointly optimized through collaborative training> Experiments on BigVul demonstrate that DCVD consistently outperforms state-of-the-art methods on both function-level detection and statement-level localization> Specifically, the framework leads on statement-level classification under both the Two-Phase and One-Phase protocols, showing improvements in MCC and F1 scores across these settings> Furthermore, DCVD achieves the best performance on ranking metrics like Top-1, MFR, and MAR compared to baselines> The ablation study confirms that removing components like Fusion or Multi-Task leads to significant performance drops, validating the necessity of deep cross-modal fusion and explicit statement-level supervision for achieving robust joint detection and localization>

Conclusion

DCVD is a unified framework that synergistically integrates multi-source information extraction, cross-modal feature fusion, and multi-granularity supervised learning for joint vulnerability detection and localization> The core insights underpinning DCVD are deep cross-modal fusion between structural and semantic representations, and explicit supervision at both the function and statement levels> Comprehensive evaluations consistently validate the effectiveness of this design choice on BigVul.

<ref:2605.11015#pg10>

<ref:2605.11015#pg10>

<ref:2605.11015#pg10>

<ref:2605.

Improvements for AI systems

  1. Bold header: Dual-Channel Encoder for Multi-Dimensional Feature Extraction

This encoder combines GAT-based control dependency feature extraction with LLM-enhanced semantic feature extraction in a parallel architecture, effectively capturing multi-dimensional information closely related to vulnerabilities from both structural and semantic perspectives.

  1. Bold header: Cross-Modal Fusion Module for Representation Alignment

The system uses this module, which employs contrastive learning and bidirectional cross-attention, to bridge the representation gap between graph-structural and textualsemantic features by aligning modalities into a shared space before enabling deep interaction and mutual enhancement.

  1. Bold header: Multi-Granularity Supervisor for Collaborative Optimization

This component is designed to address supervision asymmetry by incorporating both function-level and statement-level signals, allowing the system to achieve collaborative detection and localization without relying on indirect attention-based attribution.

Sources

Related papers