Understanding Backdoor Vulnerabilities in Vertical Federated Learning: The Gap Between Research and Practice

arXiv:2608.12962 · cs.LG, cs.CR · Submitted 2026-08-13 · Read on arXiv

Ziqi Zhao, Jialin Lu, Junjie Shan, Junyuan Zhang, Shuya Yang, Ka-Ho Chow

The University of Hong Kong

cs.LG, cs.CR

Submitted: 2026-08-13

Updated: 2026-08-14

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: This paper, "Understanding Backdoor Vulnerabilities in Vertical Federated Learning: The Gap Between Research and Practice" by Ziqi Zhao, Jialin Lu, Junjie Shan, Junyuan Zhang, Shuya Yang, and Ka-Ho

Terminology

Summary

This paper, Understanding Backdoor Vulnerabilities in Vertical Federated Learning: The Gap Between Research and Practice by Ziqi Zhao, Jialin Lu, Junjie Shan, Junyuan Zhang, Shuya Yang, and Ka-Ho Chow from the School of Computing and Data Science at The University of Hong Kong, presents a systematic, practice-oriented study of backdoor vulnerabilities in Vertical Federated Learning (VFL).

Vertical Federated Learning (VFL) enables organizations holding complementary features of shared entities to collaborate and train models. In this setting, the initiator (active party) can withhold information about the learning task, while other contributors (passive parties) participate without exposing their local datasets, creating an asymmetric information structure. The paper notes that this asymmetry is a double-edged sword - while it aligns with growing privacy demands, it also creates unique security risks, particularly backdoor attacks.

The authors argue that the current understanding of backdoor vulnerabilities in VFL is fragile and identify a fundamental gap between research and practice. They attribute this gap to misalignment along two dimensions: methodological design and evaluation design.

The paper identifies that existing attacks rely on unrealistic assumptions regarding threat models. Specifically:

  • Target-Class Label (Ltarget): All attacks require the adversary to identify or infer training samples belonging to the target class. The authors state this is unlikely to be met because the active party is not obligated to disclose task-specific information.

  • Full Label Space (Lfull): Some attacks assume knowledge of the entire label space, which is even more unlikely to be accessible than the target-class label alone.

  • Top Model Posterior (P): Some attacks assume the adversary can observe the top model's posterior during inference, which is in direct conflict with the VFL setting where the passive parties have no access to the top model.

The paper defines a practical threat model: "The attacker possessing the capabilities as a passive party should achieve the attack goal without prior knowledge of the ML task, such as class semantics, target-class labels, full label space, or any other task-related information. Under this threat model, no existing attacks remain feasible."

The paper outlines a practical backdoor workflow: "First, label acquisition should aim to cluster training samples of the same class without access to task-related knowledge. Second, each cluster should be assigned a pseudo-label and a dedicated trigger pattern, and the class semantics of each pseudo-label should be inferred after deployment by sending trigger-injected queries and observing responses. Third, backdoor learning should be formulated as a multi-target instance while balancing effectiveness across targets."

For defenses, the paper identifies that VFL-specific defenses require prior knowledge including clean reference embeddings and knowledge of the adversarial environment such as the number of attackers and target classes. The practical threat model for defenses states: "The defender possessing the capabilities as an active party should achieve the defense goal without clean reference embeddings and the adversarial environment. Under this threat model, no existing VFL-specific defenses remain practical, as they all explicitly or implicitly rely on clean reference embeddings."

The paper identifies that current evaluation practices are fundamentally flawed because:

  • Most works compare against only a few baselines

  • Existing evaluations adopt heterogeneous configurations (different model architectures and hyperparameters)

  • Existing studies rely on artificial datasets that do not reflect real-world VFL deployments

  • Evaluations are fragmented, focusing on limited metrics while overlooking critical aspects

To address these gaps, the authors introduce BVBench, a backdoor-focused benchmark designed to enable fair, practical, and comprehensive evaluation of VFL vulnerabilities, preloaded with state-of-the-art attacks and defenses. BVBench includes:

  1. Unified VFL Engine: standardizes model architectures, training configurations, and communication protocols while strictly enforcing information asymmetry

  2. Realistic VFL Datasets: Includes Satellite, KUHAR, PTB-XL, Vehicle, and NUSWIDE datasets that capture natural cross-party data characteristics

  3. State-of-the-art Attacks and Defenses: Nine attacks and eight defenses, all implemented in PyTorch with over 20,000 lines of code

The benchmark provides standardized evaluation recipes covering efficacy, dependency, stability, robustness, stealthiness, and overhead for attacks; and efficacy, dependency, stability, robustness, and overhead for defenses.

The paper finds that current VFL backdoor attacks are considerably less effective than previously implied. Key results include:

  • Although all attacks preserve MTA on CIFAR-10, they can incur substantial utility degradation on realistic datasets, reducing MTA by as much as 13.39% on Satellite, 50.50% on KUHAR, and 17.78% on PTB-XL.

  • Despite prior studies often reporting ASR exceeding 80% on CIFAR-10, BadVFL, LFBA, and VILLAIN achieve less than 30% ASR under our evaluation

  • No existing attack demonstrates consistently strong performance across datasets. The gap between the best and worst cases can be substantial, with ASR varying by as much as 92.39% for BadVFL*

The paper reveals that Almost all attacks achieve their highest ASR when targeting Class 2, which corresponds to the majority class. In multi-target settings, attacks appear to preferentially optimize shortcuts associated with majority classes when trade-offs arise. Consequently, ASR on minority classes can deteriorate drastically.

Using LFBA as a case study, the paper finds that reducing label quality by only 20% can decrease ASR by more than 25% and that label knowledge is neither sufficient nor uniformly obtainable in practice.

The paper shows that attacks can be highly sensitive to randomness: An attack may perform well during one epoch but fail in the next, or succeed under one random seed while collapsing under another. For example, BAEVFL's ASR on Vehicle drops from 100% to 0% within two consecutive epochs and recovers to 100% only two epochs later.

The paper finds that Embedding-space attacks consistently outperform input-space attacks regardless of the number of parties. Additionally, "coordinated attackers targeting the same class substantially improve ASR relative to the single-attacker baseline. In contrast, independent attackers pursuing different targets interfere with one another and significantly degrade attack effectiveness."

The paper reveals that current VFL backdoor attacks are not stealthy from the defender's perspective. Effective attacks tend to be highly anomalous, whereas stealthier attacks are often ineffective. The ROC-AUC between malicious and benign samples approaches 1.0, indicating near-perfect detectability when ASR is high.

The paper finds that No existing defense can simultaneously preserve benign utility, suppress attacks, and recover correct predictions on attacked samples. Key findings include:

  • Existing defenses exhibit a serious utility-security trade-off. Methods that effectively suppress attacks often incur substantial degradation in clean performance

  • No defense achieves the objective of utility recovery approaching MTA, including methods explicitly designed for recovery

  • For example, VFLMonitor reduces ASR from 73.68% to 29.03% on Vehicle, yet its UR reaches only 31.79%, far below its MTA of 62.52%

The paper finds that current defenses provide uneven protection and may leave certain classes substantially more vulnerable than others. For instance, VFLIP reduces the ASR of Classes 0, 1, and 3 to below 10%, yet strengthens the attack on the majority class (Class 2).

The paper reveals that existing defenses provide only limited protection against colluding attackers and that defenses can inadvertently strengthen attacks by redistributing protection unevenly across targets.

The paper concludes that current understanding of VFL backdoor vulnerability is incomplete and unreliable, and should be calibrated under the practical threat model and evaluation recipe to guide future research. The authors state that "On the attack side, no existing attack remains practical, because they all rely on inaccessible task-related knowledge. Even with such knowledge, their effectiveness is still limited by strong dependency, instability, multi-target incompatibility, poor stealthiness, and sensitivity to the VFL environment."

The paper emphasizes that This does not mean VFL systems are already well protected against vulnerability. Current defenses rely on unreasonable knowledge and fail to consistently achieve all three defense objectives.

The authors note limitations of their study: We focus on neural-network-based VFL and single-label classification, while real-world VFL may involve other tasks or protocols. They also clarify that this work aims to recalibrate the understanding of practical VFL backdoor vulnerabilities rather than propose a new practical attack or defense.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems:

1. Task-Agnostic Label Acquisition Module

  • I can implement a clustering-based label inference system that groups training samples by class without requiring any task-specific knowledge (class semantics, label space, or target labels)

  • The system can assign pseudo-labels to clusters and infer actual class semantics post-deployment by sending trigger-injected queries and observing responses

  • This enables backdoor attacks to operate under realistic VFL threat models where the active party withholds task information

2. Multi-Target Backdoor Optimization

  • I can formulate backdoor learning as a multi-target optimization problem that balances effectiveness across all target classes simultaneously

  • The system can dynamically allocate trigger patterns and learning weights to prevent majority-class bias, ensuring minority classes are not sacrificed

  • This addresses the observed failure where attacks preferentially optimize shortcuts for majority classes, causing minority-class ASR to deteriorate drastically

3. Stability-Aware Attack Scheduling

  • I can implement epoch-level and seed-level stability monitoring that detects performance collapse (e.g., ASR dropping from 100% to 0% within consecutive epochs)

  • The system can automatically trigger re-initialization or adaptive learning-rate adjustments when instability is detected

  • This mitigates the randomness sensitivity observed in attacks like BAEVFL on Vehicle dataset

4. Stealthiness-Constrained Attack Generation

  • I can incorporate a stealthiness constraint directly into the attack objective function, penalizing anomalous embedding patterns that make malicious samples distinguishable (ROC-AUC approaching 1.0)

  • The system can generate attacks that maintain high ASR while keeping malicious samples statistically indistinguishable from benign ones in the embedding space

  • This addresses the fundamental trade-off where effective attacks are highly detectable

5. Defense Without Clean Reference Embeddings

  • I can implement defense mechanisms that operate without clean reference embeddings or knowledge of adversarial environments (number of attackers, target classes)

  • The system can use online anomaly detection based on embedding distribution statistics computed from live traffic, rather than pre-collected clean baselines

  • This makes defenses practical under the realistic threat model where defenders lack prior knowledge

6. Utility-Security Balanced Defense

  • I can design a defense that jointly optimizes three objectives: preserving benign utility (MTA), suppressing attack success rate (ASR), and recovering correct predictions on attacked samples (UR)

  • The system can dynamically adjust protection strength per class, avoiding the observed failure where defenses strengthen attacks on majority classes while protecting minority classes

  • This addresses the serious utility-security trade-off where current defenses either suppress attacks but degrade clean performance, or preserve utility but fail to recover attacked samples

7. Collusion-Resistant Defense

  • I can implement defense mechanisms that remain effective against coordinated multi-attacker scenarios where attackers collude on the same target class

  • The system can detect and neutralize coordinated embedding manipulations by cross-referencing multiple passive parties' contributions and identifying correlated anomalies

  • This addresses the finding that current defenses provide limited protection against colluding attackers

8. Fair Evaluation and Benchmarking Framework

  • I can build an evaluation harness that standardizes model architectures, training configurations, and communication protocols across all attack-defense combinations

  • The system can automatically report standardized metrics (efficacy, dependency, stability, robustness, stealthiness, overhead) using realistic datasets like Satellite, KUHAR, PTB-XL, Vehicle, and NUSWIDE

  • This enables fair comparison and prevents misleading conclusions from heterogeneous evaluation configurations

The improved AI system can:

  • Conduct realistic VFL backdoor attacks without any task-related knowledge, using clustering and pseudo-label inference to achieve attack goals under practical constraints

  • Generate stable, stealthy, multi-target attacks that maintain consistent performance across epochs and seeds while remaining undetectable by defenders

  • Deploy practical defenses that work without clean reference data and provide balanced protection across all classes without utility degradation

  • Resist colluding attackers by detecting correlated anomalies across multiple parties

  • Provide reliable security assessments through standardized benchmarking, enabling researchers and practitioners to make informed decisions about VFL security investments

  • Recover correct predictions on attacked samples while simultaneously preserving benign utility and suppressing attacks, achieving all three defense objectives simultaneously

Abstract

Vertical Federated Learning (VFL) enables organizations holding complementary features of shared entities to collaborate and train models. In this setting, the initiator can withhold information about the learning task, while other contributors participate without exposing their local datasets, creating an asymmetric information structure aligned with growing privacy demands. However, this asymmetry is a double-edged sword. Among various threats, backdoor attacks are particularly concerning because VFL not only enables malicious contributors to poison the model during training, but also allows them to activate the backdoor at inference time to manipulate predictions. Although prior work has reported near-perfect attack success rates and proposed effective defenses, we find that most findings fail to hold under realistic conditions, exposing a fundamental gap between research and practice. In this paper, we present a systematic, practice-oriented study of backdoor vulnerabilities in VFL, revealing this gap in both methodological design and evaluation practices. We show that existing approaches overlook key practical constraints and therefore rely on unrealistic prior knowledge. Furthermore, these limitations have remained hidden due to poorly designed evaluation practices in the literature. To bridge this gap, we redefine threat models under realistic constraints, propose practical backdoor workflows, and introduce BVBench, a backdoor-centric benchmark that enables fair, practical, and comprehensive evaluation, preloaded with state-of-the-art baselines. BVBench provides strong evidence of the fragility of the current understanding of VFL backdoor risks and establishes a foundation for steering research toward uncovering practical vulnerabilities and developing more meaningful defenses.

Sources

Related papers