Analysis of Federated Aggregation under Model Poisoning and Backdoor Attacks: A Reconstructed Cross-Dataset and Cross-Architecture Benchmark
Soumya Mazumdar, Vineet Kumar Rakesh, Tapas Samanta
Gargi Memorial Institute of Technology · Maulana Abul Kalam Azad University of Technology · Variable Energy Cyclotron Centre · Homi Bhabha National Institute
cs.LG, cs.CR, cs.CV
Submitted: 2026-08-11
Updated: 2026-08-13
Comments: 30 pages, 7 main figures, 7 main tables; includes 11 pages of Supplementary Information with 14 supplementary figures. Code and reproducibility resources: https://github.com/mazumdarsoumya/RobustFL-Bench
Code: https://github.com/mazumdarsoumya/RobustFL-Bench
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 75/100
The gist: The paper presents a reconstructed comparative benchmark and evidence audit of federated aggregation methods under model poisoning and backdoor attacks, rather than a new-algorithm superiority study.
Terminology
Summary
The paper presents a reconstructed comparative benchmark and evidence audit of federated aggregation methods under model poisoning and backdoor attacks, rather than a new-algorithm superiority study. The study analyzes a fixed 500-cell seed-1 evaluation matrix spanning five aggregation methods (FedAvg, Trimmed Mean, Krum, FLTrust, FedPARETO), five datasets (GTSRB, SVHN, MNIST, CIFAR-10, CIFAR-100), five architectures (SimpleCNN, ResNet-18, MobileNetV3-Small, EfficientNet-B0, ShuffleNetV2), and four recorded conditions (Clean, Sign flipping, Gaussian, BadNets). Each cell contains a numerical performance summary.
The reconstruction resolved an earlier incomplete 464/500 matrix by recovering numerical run-summary values for all 500 canonical identities. Successful execution logs were identified for 454 original runs and 36 repaired or rerun executions, whereas 10 clean SVHN cells (MobileNetV3-Small and ShuffleNetV2 across all five methods) were supported by summary-only provenance. The workbook cross-check is exact after a documented rounding tolerance of 5 × 10−5.
Key empirical findings include: Trimmed Mean achieved the highest clean macro-mean accuracy (76.02%) and the lowest mean within-task rank (1.70), followed by FedAvg (70.92%, rank 1.82). Krum attained the highest recorded accuracy under both sign-flipping (mean accuracy 63.91%, mean rank 1.32, 19 configuration wins) and Gaussian configurations (mean accuracy 64.45%, mean rank 1.08, 24 wins). The BadNets ordinary-accuracy ordering returns to a pattern closer to Clean, with Trimmed Mean having the highest macro-mean ordinary accuracy (75.17%) and FedAvg and Trimmed Mean tied at mean rank 1.80.
The mean matched attack-minus-clean change under sign flipping is −65.32 pp for FedAvg, −34.54 pp for Trimmed Mean, +1.06 pp for Krum, −16.41 pp for FLTrust, and −36.30 pp for FedPARETO. The comparable changes in Gaussian are −59.04, −52.88, +1.60, −16.07, and −52.62 pp. The tiny positive Krum means are not understood as attack-induced improvement; rather, the supporting finding is that Krum's reported attacked accuracies are similar to its clean accuracies in the seed-1 matrix.
These relative rankings remained unchanged when analysis was restricted to 21 task pairs for which original successful logs were available for every method–condition combination. The provenance sensitivity analysis shows that the Krum-first ordering under Sign flipping and Gaussian is retained in both provenance-sensitive subsets. Krum's mean matched attack-minus-clean change is +0.84 pp (Sign flipping) and +1.58 pp (Gaussian) in the 23-task set, and +0.83 pp and +1.42 pp in the 21-task all-original set, compared with +1.06 pp and +1.60 pp in the full matrix.
The audit of the supplied BadNets metric implementation established that every test input is triggered prior to target-label counting; consequently, the retained metric represents Triggered Target-Label Rate (TTLR) rather than a conventional target-excluding attack success rate. Under the supplied audited metric semantics, mean TTLR is 72.68% for FedAvg, 69.70% for Trimmed Mean, 60.54% for FedPARETO, 30.61% for Krum, and 24.41% for FLTrust. The low mean TTLR of FLTrust occurs together with low ordinary accuracy, illustrating why TTLR should not be translated directly into a backdoor-defense ranking.
An audit of the supplied FedPARETO scaffold further identified a pathway in which predictive summaries may characterise an uncorrupted local model while the aggregation weight is applied to a separately corrupted update, introducing a potential discrepancy between reported predictive outcomes and the updates used for aggregation. In the audited sign-flipping/Gaussian client path, local training first creates an uncorrupted model and delta. The delta may then be modified by the attack, while local metrics and anchor predictions are subsequently computed from the still-uncorrupted local model object. The executable scaffold constructs four raw signals: anchor accuracy, negative anchor ECE, low-local-accuracy priority (1 − local client accuracy), and a robustness score combining root-direction similarity and temporal reliability. The executable approach is a Pareto-inspired dominance-count/scalarization heuristic; it does not implement the conceptual draft's weighted-Tchebycheff optimization, explicit l1 trust region, reliability-dependent weight caps, benign-subspace singular-vector model, or formal constrained Pareto solver.
The canonical matrix contains a single identified seed for each cell, and the exact attack and configuration lineage is incomplete. No p-values, t-tests, ANOVA, seed-level standard errors, or seed-level confidence intervals are reported for the canonical matrix because there is one identified seed per cell. The reported findings should be interpreted as descriptive comparisons within the recorded configurations and should not be construed as statistical estimates or universal claims regarding robustness. The study does not include adaptive Krum-aware optimization, model replacement, attack-strength sweeps, or controlled heterogeneity sweeps. The strongest defensible conclusion is that Krum recorded the best final-accuracy ordering among the five evaluated methods under the retained Sign-flipping and Gaussian condition labels.
Improvements for AI systems
Improvements to AI Systems:
-
Attack-Aware Aggregation Selection: Implement a dynamic aggregation-method selector that switches between Trimmed Mean (for clean/high-accuracy tasks) and Krum (for sign-flipping or Gaussian attack conditions) based on real-time anomaly detection, improving robustness without sacrificing clean performance.
-
Provenance-Aware Evaluation Pipeline: Build an AI system that automatically audits experimental matrices for missing cells, rounding errors, and provenance gaps (e.g., summary-only vs. log-backed results), flagging untrustworthy comparisons before drawing conclusions—reducing false confidence in benchmark claims.
-
Metric-Semantics Disambiguation: Enhance model-evaluation frameworks to automatically detect and correct metric misinterpretations (e.g., distinguishing Triggered Target-Label Rate from conventional attack success rate), preventing AI systems from overestimating backdoor defense efficacy.
-
Scaffold-Integrity Verification: Integrate a runtime checker that verifies whether aggregation weights are computed from the same model objects used for predictive summaries, detecting discrepancies between reported performance and actual update corruption—improving trust in federated learning logs.
-
Heterogeneity-Robust Rank Aggregation: Develop a meta-learner that combines multiple aggregation methods (FedAvg, Trimmed Mean, Krum, FLTrust, FedPARETO) using within-task rank-based weighting, rather than raw accuracy, to produce stable rankings under varied architectures and datasets—improving generalization across unseen tasks.
-
Seed-Uncertainty Quantification: Add a module that, when only single-seed results exist, automatically generates conservative confidence intervals using cross-task variance and provenance sensitivity subsets, preventing overclaiming of statistical significance in federated robustness studies.
What the Improved AI System Can Do:
-
Automatically select the most robust aggregation method per attack scenario (e.g., Krum under sign-flipping) while maintaining high clean accuracy via Trimmed Mean, improving real-world federated deployment resilience.
-
Detect and repair incomplete or inconsistent benchmark matrices, ensuring reliable comparisons across methods, datasets, and architectures without manual auditing.
-
Correctly report backdoor attack metrics (TTLR vs. standard ASR) to avoid misleading defense rankings, enabling safer AI deployment in security-sensitive applications.
-
Verify that federated learning logs are internally consistent (weights match model states), reducing vulnerability to silent corruption or adversarial manipulation of reported results.
-
Provide uncertainty-aware robustness rankings that account for limited seeds and provenance gaps, guiding researchers toward more defensible conclusions and better resource allocation for future experiments.
Abstract
Robust comparisons of federated aggregation methods require joint consideration of predictive performance, threat definitions, metric semantics, and execution provenance. A 500-cell seed-1 evaluation matrix was reconstructed across five aggregation methods, five datasets, five architectures, and four recorded conditions: clean, sign-flipping, Gaussian, and BadNets. Successful execution logs were identified for 454 original runs and 36 repaired or rerun executions, whereas 10 clean SVHN cells were supported by summary-only provenance. Trimmed Mean achieved the highest clean macro-mean accuracy (76.02%) and the lowest mean within-task rank (1.70). Krum attained the highest recorded accuracy under both sign-flipping and Gaussian configurations. These relative rankings remained unchanged when analysis was restricted to 21 task pairs for which original successful logs were available for every method-condition combination. Audit of the supplied BadNets metric implementation established that every test input is triggered prior to target-label counting; consequently, the retained metric represents Triggered Target-Label Rate (TTLR) rather than a conventional target-excluding attack success rate. An audit of the supplied FedPARETO scaffold further identified a pathway in which predictive summaries may characterize an uncorrupted local model while the aggregation weight is applied to a separately corrupted update, introducing a potential discrepancy between reported predictive outcomes and the updates used for aggregation. The canonical matrix contains a single identified seed for each cell, and exact attack and configuration lineage is incomplete. Accordingly, the findings should be interpreted as descriptive comparisons within the recorded configurations and not as statistical estimates or universal claims regarding robustness.
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks