Benchmarking Cyberattack Detection in Electric Vehicle Charging Infrastructure with Benign User Updates
Hannan Chen, Roshni Anna Jacob, Jie Zhang
The University of Texas at Dallas
cs.CR, cs.LG, cs.SY, eess.SY
Submitted: 2026-08-11
Updated: 2026-08-13
Code: https://github.com/Hannan-Chen/Cyberattack-Detection-in-EV-Charging-Systems
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: This paper develops a leakage-controlled session-level benchmark for cyberattack detection in electric vehicle (EV) charging infrastructure that preserves the ordered inputs of real Adaptive Charging
Terminology
Summary
This paper develops a leakage-controlled session-level benchmark for cyberattack detection in electric vehicle (EV) charging infrastructure that preserves the ordered inputs of real Adaptive Charging Network (ACN) sessions and models legitimate revisions as normal behavior. The study uses 1,213 valid Caltech ACN sessions, resulting in 1,264 normal states comprising 1,213 activations and 51 updates (13 demand-only, 14 departure-only, and 24 joint updates). The paper states: Charging manipulation attacks can exploit the same interface and variables; therefore, detecting a request change alone does not establish malicious intent.
The benchmark applies six physically motivated charging manipulation attacks (CMAs) and six coordinated variants to eligible request states. The six CMAs are: Type-1 (demand increase), Type-2 (earlier departure), Type-3 (later start), Type-4 (demand + earlier departure), Type-5 (demand + later start), and Type-6 (demand + both times). The attack charging rate is ACRatt = Catt Chav = 0.15(30) = 4.5 kW. A fixed pool keeps each generated attack in its source session's split, with CV attacks originating only from validation sources and final attacks only from final-test sources.
The paper compares 22 profile-only, transition-aware, and context-stratified model families under common source-grouped folds, attack data, and operating constraints, evaluating 89 configurations in 445 fold tasks. The proposed Dual-Branch Masked-Autoencoder (Masked-AE) Transition Boost model evaluates whether the current request is normal and whether its producing transition resembles an observed benign update.
Its state branch combines masked reconstruction with a radial-basis-function one-class support boundary, while its transition branch combines masked reconstruction with shrinkage covariance distance.
The proposed model architecture uses state widths (64, 32), latent dimension 4, masking probability ρs = 0.15, and an RBF OCSVM with ν = 0.05 and γ = scale; its transition AE uses widths (12, 6), latent dimension 2, masking probability ρt = 0.25, Ledoit–Wolf distance, and transition weight ωtr = 0.15. The score fusion is: sdual(n, l) = sstate(n, l) + ωtr In,l chg str(n, l), where the gate In,l chg = 0 for activation and one for an observed update or generated attack.
Source-grouped five-fold cross-validation selects complete configurations under explicit overall-normal and benign-update acceptance constraints (TNRall ≥ 0.95, TNRupdate ≥ 0.90). The selection objective is Jcv(c) = F1macro(c) − 0.5 SDfold[F1macro(c)]. Disjoint normal data then calibrate the final threshold before one test evaluation. The selected model has Jcv = 0.671, exceeding Profile-Anchor Transition Boost's 0.653.
The formal model achieves F1 = 0.823 on 1,505 final attacks (with TN=240, FP=12, FN=445, TP=1060), accepts all nine test updates (update TNR = 1.000), and gives comparable recall for activation-origin attacks (0.705) and update-origin attacks (0.692). The paper states: The developed dual-branch model provides the strongest robust validation performance while detecting malicious request manipulations without learning to reject legitimate user choices.
Key findings include: (1) recorded user updates must be modeled as legitimate transitions; (2) the state and transition branches are complementary; (3) attack physics determines separability, with Type-1-ta, Type-2-ta, and Type-4-ta being difficult because their per-profile perturbations shrink near requested departure, while Types 5-ta and 6-ta reach recalls 0.991 and 0.990. The paper concludes: Reliable CMA detection therefore requires authorized revisions to be modeled throughout splitting, learning, selection, and calibration.
Improvements for AI systems
Improvements to AI systems:
-
Dual-branch anomaly detection with explicit transition modeling: Implement a two-branch architecture where one branch reconstructs the current state (using masked autoencoders with one-class boundaries) and a second branch reconstructs the state transition (using shrinkage covariance distance). This allows the AI to distinguish between malicious request changes and legitimate user revisions by explicitly modeling the process of change, not just the resulting state.
-
Leakage-controlled evaluation protocol: Adopt source-grouped cross-validation where attack data are generated only from the same session split as their origin, preventing data leakage between training and test sets. This ensures the AI system’s performance metrics reflect real-world generalization, not inflated by overlapping session data.
-
Constraint-aware model selection with dual acceptance criteria: During hyperparameter selection, enforce explicit constraints on overall normal acceptance (TNR ≥ 0.95) and benign-update acceptance (TNR ≥ 0.90), and use an objective that penalizes fold variance (F1 macro − 0.5 × SD across folds). This yields models that are robust across session groups and do not overfit to specific attack patterns.
-
Physics-informed attack generation for training: Generate attack samples using physically motivated charging manipulation types (demand increase, earlier departure, later start, and combinations) with realistic attack rates (e.g., 4.5 kW). This allows the AI to learn attack signatures that are consistent with actual EV charging constraints, improving detection of coordinated and subtle manipulations.
-
Gated score fusion with transition weighting: Use a gate that activates the transition branch only for observed updates or generated attacks (not for initial activations), and fuse state and transition scores with a small transition weight (ω = 0.15). This prevents the AI from over-penalizing normal session starts while still catching malicious revisions.
-
Calibration on disjoint normal data: After model selection, calibrate the final decision threshold using a separate set of normal (non-attack) sessions that were never used in training or validation. This ensures the AI’s false-positive rate is tuned to real-world benign traffic, not just the validation fold.
What the improved AI system can do:
-
Detect charging manipulation attacks in EV infrastructure with high recall (F1 = 0.823) while accepting all legitimate user updates (100% update TNR), meaning it will not block or flag normal user changes to their charging sessions.
-
Distinguish between benign revisions (e.g., user extends departure time) and malicious manipulations (e.g., attacker increases demand) even when both alter the same request variables, by learning the statistical signature of legitimate transitions.
-
Generalize across different sessions and time periods without data leakage, providing reliable performance on unseen data.
-
Achieve near-perfect detection (recall > 0.99) for attacks that significantly alter demand or timing (Types 5 and 6), while maintaining robustness against harder-to-detect attacks that shrink perturbations near departure.
-
Provide a benchmark-driven, reproducible framework for evaluating any anomaly detection model on session-level cyberattack data, with explicit constraints on normal and update acceptance rates.
Sources
- GDGU: A Gradient Difference-based Graph Unlearning Method for Cyberattack Localization in Electric Vehicle Charging Networks
- Gaussian Error Linear Units (GELUs)
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs