Cross-Corpus Evaluation of Generalizable Vulnerability Detection in IoT Firmware

arXiv:2608.11492 · cs.CR, cs.LG · Submitted 2026-08-15 · Read on arXiv

Sadib Hassan Rumman, Md. Shariful Islam, Md Rayhanur Rahman

University of Dhaka · The University of Alabama

cs.CR, cs.LG

Submitted: 2026-08-15

Updated: 2026-08-18

Comments: 6 pages, 1 Figure, 2 Tables

Code: https://github.com/rumman9799/IoTVulBench

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

The gist: This paper introduces IoTVulBench, a human-verified benchmark for cross-corpus IoT firmware vulnerability detection, and uses it to systematically evaluate how training-data source, model

Terminology

Summary

This paper introduces IoTVulBench, a human-verified benchmark for cross-corpus IoT firmware vulnerability detection, and uses it to systematically evaluate how training-data source, model architecture, tuning method, and curriculum design affect detection accuracy, efficiency, and robustness. The study addresses three research questions: (RQ1) whether domain-matched, human-verified IoT firmware training data outperforms general-purpose vulnerability sources and whether staged curriculum learning improves cross-corpus generalization further; (RQ2) how architecture, tuning method, and model scale affect the accuracy-efficiency trade-off, and whether ensemble or distillation strategies improve on single-model deployment; and (RQ3) how robust trained detectors are to lexical and compiler-level perturbation, distribution shift, and real-world temporal conditions, and how faithful their explanations are to true vulnerability indicators.

IoTVulBench-Core was constructed from GitHub embedded and IoT firmware repositories and verified by three independent experts rather than automated labeling, with inter-annotator agreement measured using Fleiss’ kappa for each CWE category and disagreements resolved through a documented adjudication process. A fixed, contamination-screened held-out target was partitioned from the verified corpus and excluded from all training conditions. Four primary training sources distinguished by label provenance were compared—IoTVulBench, D2A, SARD-Embedded, and PrimeVul—and all datasets were harmonized and randomly undersampled to match IoTVulBench. Training sources were screened against the held-out target using 8-gram character-overlap matching and repository-level fork detection, with flagged samples removed before training.

Five architectures spanning three classes were evaluated: UniXcoder-base (encoder class), CodeLlama-7B and Qwen2.5-Coder-7B (small decoder class), and Qwen2.5-Coder-32B and DeepSeek-Coder-V2-16B (mid-scale decoder class). Decoders used LoRA/QLoRA, whereas the encoder was fully fine-tuned. Each training source was evaluated under size-matched and full-size settings using single-source, pooled, and staged curricula, with top configurations repeated across multiple seeds. Performance was compared against a static analyzer, a prior encoder-based detector, and a frontier LLM, using MCC as the primary metric, with precision, recall, F1, false-negative rate at a fixed false-positive rate, calibration, and severity-weighted MCC also reported.

For RQ1, under matched conditions, models trained with IoTVulBench achieved the highest MCC (0.58) of any single source, exceeding PrimeVul (0.44), D2A (0.39), and SARD-Embedded (0.29). The source ranking held under both size-matched and full-size training, indicating the advantage reflects label quality and domain match rather than data volume. Staged curriculum learning (pretraining on pooled generic sources, then fine-tuning on IoTVulBench) raised MCC to 0.69, exceeding both single-source (0.58) and naive pooling (0.53), with precision and recall both improving under staging. The staged IoTVulBench model’s advantage over PrimeVul, the strongest single-source training dataset, remained significant after correction, confirming the domain-match effect holds against the best available alternative.

For RQ2, architecture performance depended on the training source, with no model consistently achieving the highest MCC (Spearman ρ = 0.20, p = 0.42). At matched scale (≤7B), LoRA-tuned decoders outperformed fully fine-tuned encoders. The staged Qwen2.5-Coder-7B achieved an MCC of 0.69, increasing to 0.73 with a diversity-optimized ensemble, surpassing all baselines by at least ∆MCC ≥ 0.42. QLoRA-native INT4 quantization retained 95% of full-precision performance, while a distilled 350M model preserved 91% of the teacher’s MCC. A recall-optimized cascade achieved MCC 0.70 using an average of 1.1B active parameters. The distilled model and cascade provided the best accuracy-efficiency trade-off, whereas the ensemble delivered the highest detection performance.

For RQ3, the proposed model retained 86% and 82% of unperturbed MCC under identifier renaming and compiler/optimization shifts, respectively, indicating structural rather than lexical dependence. At a 0.5% false-positive-rate tolerance, it missed 21% of vulnerabilities versus 71% for the strongest comparator, with strong calibration (Brier score = 0.09). Explanations overlapped vulnerable lines in 78% of correct detections, with 71% causally validated and 0.74 localization accuracy. Shift-robust conformal prediction maintained coverage within 1.2 points of 95% versus 8–11 points of under-coverage for standard methods, while refinement improved MCC by 0.02 per cycle without detected drift. Extended validation achieved MCC 0.65 on screened post-cutoff CVEs and 0.64 under federated training, within 0.05 of centralized training; hardware-in-the-loop testing confirmed 29/40 detections as exploitable. All pre-registered comparisons remained significant after Benjamini–Hochberg correction (paired Wilcoxon, N = 96): IoTVulBench versus generic sources (∆MCC = +0.14–+0.29, p < 0.02), staged versus single-source curricula (∆MCC = +0.11, p = 0.008), and the proposed configuration versus the strongest static analyzer (∆MCC = +0.42, p < 0.001).

The paper concludes that firmware vulnerability detection depended more on label provenance and curriculum design than model scale: a staged-trained 7B model outperformed a 32B model trained on generic data, and no architecture generalized best across all training sources. Robustness, calibration, and largely faithful explanations suggest that predictions relied primarily on structural reasoning, although the 71% causal correctness rate indicates room for improvement. Temporal, federated, and hardware-in-the-loop evaluations further supported generalizability, with hardware evidence remaining limited. Future work will expand robustness testing, hardware validation, explanation faithfulness, rare-CWE coverage, federated evaluation, and patch generation, and will incorporate full-scale versions of external datasets alongside an enlarged IoTVulBench corpus.

Improvements for AI systems

Improvements to AI Systems:

  1. Domain-matched, human-verified training data over generic sources: Train vulnerability detectors on firmware-specific, expert-annotated datasets (like IoTVulBench) rather than general-purpose vulnerability corpora. This improves detection accuracy (MCC 0.58 vs. 0.29–0.44 for generic sources) and reduces false positives, enabling more reliable identification of real IoT firmware flaws.

  2. Staged curriculum learning for cross-corpus generalization: Implement a two-stage training pipeline—first pretrain on pooled generic vulnerability data, then fine-tune on domain-matched firmware data. This raises detection MCC from 0.58 (single-source) to 0.69, improving both precision and recall, allowing the system to generalize better to unseen firmware codebases.

  3. Architecture-agnostic selection based on training source: Avoid assuming a single best model; instead, evaluate multiple architectures (encoder, small decoder, mid-scale decoder) per dataset. Use LoRA/QLoRA for decoders and full fine-tuning for encoders, as performance varies by source (Spearman ρ = 0.20). This enables adaptive model choice for optimal accuracy-efficiency trade-offs.

  4. Diversity-optimized ensembling for peak performance: Combine multiple staged-trained models (e.g., Qwen2.5-Coder-7B variants) with diversity-aware selection to boost MCC from 0.69 to 0.73, surpassing all baselines by ≥0.42. This improves detection of rare and complex vulnerabilities in resource-constrained IoT environments.

  5. Distillation and cascading for efficient deployment: Use knowledge distillation to compress a 7B teacher into a 350M student (retaining 91% MCC) and a recall-optimized cascade with 1.1B active parameters (MCC 0.70). This enables real-time, on-device vulnerability scanning in embedded systems with limited compute and memory.

  6. Quantization-aware fine-tuning for robustness: Apply QLoRA-native INT4 quantization during training to retain 95% of full-precision performance, reducing memory footprint and inference cost, making large-model detection feasible on edge hardware.

  7. Structural reasoning over lexical patterns: Train models to rely on code structure rather than identifier names, as evidenced by 86% and 82% MCC retention under identifier renaming and compiler/optimization shifts. This improves robustness against obfuscation and code transformations in real-world firmware.

  8. Calibrated confidence and shift-robust conformal prediction: Integrate conformal prediction that maintains 95% coverage within 1.2 points under distribution shift (vs. 8–11 points under-coverage for standard methods). This provides reliable uncertainty estimates, enabling safe deployment where false negatives are costly (e.g., missing 21% vs. 71% of vulnerabilities at 0.5% FPR).

  9. Explainability with causal validation: Generate explanations that overlap vulnerable lines in 78% of correct detections, with 71% causally validated and 0.74 localization accuracy. This allows security analysts to trust and act on detections, and enables automated patch generation by pinpointing root causes.

  10. Temporal and federated learning adaptability: Train and update models on post-cutoff CVEs (MCC 0.65) and via federated learning (MCC 0.64, within 0.05 of centralized), allowing continuous improvement across distributed firmware repositories without sharing sensitive code, and maintaining performance on newly emerging vulnerabilities.

What the improved AI system can do:

  • Detect IoT firmware vulnerabilities with high accuracy (MCC 0.69–0.73) across unseen codebases, outperforming static analyzers and frontier LLMs by large margins.

  • Operate efficiently on edge devices via distilled (350M) or cascaded (1.1B) models, enabling real-time scanning during firmware compilation or deployment.

  • Remain robust to code obfuscation, compiler changes, and temporal shifts, with calibrated confidence for risk-based prioritization.

  • Provide causally validated, line-level explanations for each detection, facilitating rapid human review and automated remediation.

  • Continuously improve via federated learning across organizations, preserving data privacy while adapting to new vulnerability patterns.

Related papers