A SISA-based Machine Unlearning Framework for Power Transformer Inter-Turn Short-Circuit Fault Localization

arXiv:2603.06962 · eess.SY, cs.LG, cs.SY · Submitted 2026-03-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "A SISA-based Machine Unlearning Framework for Power Transformer Inter-Turn Short-Circuit Fault Localization".

Rosa: —In practical data-driven applications on electrical equipment fault diagnosis, training data can be poisoned by sensor failures, which can severely degrade the performance of machine learning (ML) models.

Dev: First, who's behind it and why it matters.

Title and authors: Rosa: So, to kick things off, we're talking about the paper titled "A SISA-based Machine Unlearning Framework for Power Transformer Inter-Turn Short-Circuit Fault Localization" and who put it together. This research is focused on solving that problem of training data getting poisoned by sensor failures in electrical equipment fault diagnosis.

Dev: I see the title, and it immediately tells me this paper is tackling a specific type of ML maintenance challenge, which is making sure the diagnostic models don't get corrupted by bad input data during operation. The authors are Liu, Yan, Sun, and Zhang from the University of Texas at Dallas and Idaho National Laboratory.

Taro: I’m interested in the context here; when you look at this work alongside other papers we've seen on Agentic AI for Scalable and Robust Optical Systems Control or Topology-Aware Reinforcement Learning over Graphs for Resilient Power Distribution Networks, this paper feels very grounded in physical reliability engineering.

Rosa: That grounding is exactly what makes it interesting, Taro; they aren't just theorizing about data poisoning abstractly; they are applying a SISA framework directly to the power transformer ITSCF localization task, which is a very tangible piece of equipment.

Dev: The implication for us as engineers is that if we can build ML models that can selectively forget bad training examples without retraining the whole thing, it drastically lowers our maintenance overhead and speeds up how quickly we can update those models when real-world data quality degrades.

Taro: If this works effectively in the lab, I wonder if its implications stretch to remote monitoring systems where hardware failures are common; imagine a system that can self-heal its knowledge base without requiring a full system reboot.

Rosa: That’s the vision, Taro; imagine an AI system deployed remotely on a grid component that can detect sensor failure and immediately isolate and retrain only the affected data shard to keep making accurate decisions.

Dev: The paper proposes this as a direct solution to the difficulty of removing poisoned data after initial training because full retraining is too computationally intensive and time-consuming for industrial settings.

Taro: It's interesting how they frame it as an unlearning mechanism rather than just a simple retraining protocol; it’s about surgically removing the negative influence of sensor failures.

Rosa: Precisely, Taro; it moves us from reactive maintenance to proactive knowledge management for our AI systems in critical infrastructure.

Dev: The core idea is that partitioning the data into shards and slices allows each shard to be trained independently, which is a key mechanism they are highlighting here.

The paper's summary: Rosa: Now that we’ve talked about the setup, let’s look at what the paper actually summarizes as its main contribution regarding this SISA-based Machine Unlearning Framework for Power Transformer Inter-Turn Short-Circuit Fault Localization. Essentially, they summarize their core proposal and how it addresses the contamination problem.

Dev: The summary highlights that they propose a SISA method using a softmax probability averaging strategy to handle the aggregation of predictions from multiple shard models, which is how the final output is formed.

Taro: That aggregation step sounds like a critical piece of engineering; ensuring that combining independent model predictions results in a robust final decision, especially when some underlying data streams might be compromised.

Rosa: Right, Taro; they detail how the SISA method partitions training data into shards and slices to ensure the influence of any single data point is localized within specific models through independent training processes.

Dev: And crucially, when poisoned or contaminated data points are detected, the framework only needs to retrain those affected shards starting from the compromised slice, which is where they show it efficiently reduces computational cost compared to a full retraining effort.

Taro: That targeted retraining mechanism is what makes it powerful; you avoid the massive computational drain of re-learning everything when only a small segment of the data set has been compromised by sensor errors.

Rosa: So, in short, they demonstrate that this framework restores diagnostic accuracy while significantly reducing the required retraining time compared to doing a complete model retraining from scratch.

Dev: That's the practical takeaway: high accuracy is maintained with much faster update cycles when dealing with data poisoning caused by sensor failures in transformer fault localization.

The paper's improvements: Rosa: Let’s move on to what the authors specifically highlight as the improvements their SISA-based Machine Unlearning Framework offers over existing methods, focusing on the mechanisms they developed.

Dev: The primary improvement they point to is integrating a softmax probability averaging strategy for combining predictions from multiple shard models, which serves as their mechanism for forming an aggregated prediction probability p hat(cx) before determining the final label.

Taro: That averaging technique is smart because it smooths out the individual model predictions from each shard, making the final output less sensitive to any single compromised model that might have been influenced by poisoned data.

Rosa: Exactly; they also show how this architecture ensures that when contaminated data is found, only the affected shard needs to be retrained starting from the compromised slice, which is a major architectural improvement over methods that might attempt broader updates.

Dev: And they quantify the efficiency gains: when setting the number of shards to two, SISA unlearning reduces retraining time to two hundred twenty-one point eight seconds, and it drops further to one hundred twelve point two seconds when four shards are applied, achieving speed-ups of two point zero one times and three point nine seven times respectively in terms of retraining duration for the ITSCF localization task.

Taro: That quantifiable speed improvement is what really speaks to practical deployment; we can use that data to plan our hardware updates knowing exactly how much faster the model maintenance cycle will be when we increase the sharding strategy from two to four.

Rosa: The authors also show that this approach restores diagnostic accuracy close to full retraining, even when compared against non-SISA full retraining, showing that it doesn't sacrifice too much reliability for the speed gained.

Dev: That level of accuracy restoration is important; we need to make sure that the method doesn't introduce new failure modes where localized updates cause other parts of the model to degrade unexpectedly, which is always a concern for loop rate stability.

Conclusion: Rosa: So, to wrap up on this paper, we’ve covered how the SISA-based Machine Unlearning Framework for Power Transformer Inter-Turn Short-Circuit Fault Localization works and the specific improvements it brings. It really shows a practical path forward for managing data contamination in our ML pipelines.

Dev: In essence, we learned that by partitioning data into shards and slices, you can achieve targeted retraining on affected components without incurring the full cost of retraining the entire model every time sensor data becomes faulty.

Taro: I think this framework offers a solid methodology for handling data poisoning in sequential systems; it’s a concrete strategy for when the world misbehaves by allowing us to surgically correct the knowledge base instead of scrapping it entirely.

Rosa: It really points toward more robust, maintainable AI systems that can handle the messy reality of real-world sensor data, especially in critical areas like power transformers.

Dev: We need to keep tracking how this performs under sustained stress, because while the computational benefits are clear, we still have to ensure that the localized updates don't introduce unexpected instability into the overall system loop rate.

Taro: I think the future work should focus on testing this against more complex fault conditions that involve correlated sensor failures, moving beyond isolated noise to more systemic failures.

Rosa: Indeed, Taro; moving from simulated conditions to handling systemic failures is where we need to take this research next for real-world applicability.

Dev: Alright team, that's our discussion on "A SISA-based Machine Unlearning Framework for Power Transformer Inter-Turn Short-Circuit Fault Localization." We’re ready to move on to whatever paper comes next.

The University of Texas at Dallas · Idaho National Laboratory

eess.SY, cs.LG, cs.SY

Submitted: 2026-03-07

Updated: 2026-03-07

Journal ref: 2026 IEEE Power & Energy Society General Meeting (PESGM)

DOI: 10.1109/PESGM58988.2026.11693380

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 81/100

The gist: Abstract—In practical data-driven applications on electrical equipment fault diagnosis, training data can be poisoned by sensor failures, which can severely degrade the performance of machine

Key concepts

Data Poisoning
This occurs when training data is corrupted by sensor failures in electrical equipment. This contamination can severely degrade the performance of machine learning models used for fault diagnosis.
SISA Framework
The SISA method uses a softmax probability averaging strategy to aggregate predictions from multiple independent shard models, forming a robust final output prediction before determining the final label.
Machine Unlearning
This is the process of selectively forgetting bad training examples without retraining the entire model. The framework achieves this by only retraining affected data shards, which is more efficient than full retraining.
Shards and Slices
The paper partitions training data into smaller shards and slices. This allows each shard to be trained independently, localizing the influence of any single data point within specific models.

Terminology

Summary

Abstract—In practical data-driven applications on electrical equipment fault diagnosis, training data can be poisoned by sensor failures, which can severely degrade the performance of machine learning (ML) models. However, once the ML model has been trained, removing the influence of such harmful data is challenging, as full retraining is both computationally intensive and time-consuming. To address this challenge, this paper proposes a SISA (Sharded, Isolated, Sliced, and Aggregated)-based machine unlearning (MU) framework for power transformer inter-turn short-circuit fault (ITSCF) localization. The SISA method partitions the training data into shards and slices, ensuring that the influence of each data point is isolated within specific constituent models through independent training. When poisoned data are detected, only the affected shards are retrained, avoiding retraining the entire model from scratch. Experiments on simulated ITSCF conditions demonstrate that the proposed framework achieves almost identical diagnostic accuracy to full retraining, while reducing retraining time significantly.

The key insight of SISA is to partition the training dataset into multiple independent shards, each of which is further divided into sequential slices. For each shard, constituent models are trained incrementally on larger subsets of data, and their predictions are aggregated to form the final ensemble output. Crucially, this architecture ensures that the influence of any individual data point is localized to a specific shard and slice, rather than being diffused across the entire model. When contaminated or poisoned data are identified, only the affected shard needs to be partially retrained, starting from the compromised slice, and efficiently reducing computational cost compared to full retraining.

In this study, we present a SISA-based MU framework for ITSCF localization in power transformers considering sensor failures. The proposed framework uses a softmax-probability averaging strategy: "Let zs(x) denote the logits of the s-th model for input x, and its softmax probability for class c is ps(cx) = exp zs(x)c / (k=1 exp zs(x)k). And the aggregated prediction probability is computed as: pˆ(cx) = 1/S Σ s=1 ps(cx). The final predicted label is determined by: yˆ(x) = arg max c pˆ(cx)."

The framework is applied to Long Short-Term Memory (LSTM) models, which are adopted as the baseline ML model for ITSCF localization due to their suitability for capturing temporal dependencies in sequential data. The power transformer ITSCF dataset D contains 48 operating conditions across six fault labels—HA, HB, HC, LA, LB, and LC—corresponding to the HV and LV sides in phase A, B or C of the power transformer connected to a wind turbine. The dataset is first evenly partitioned into S shards (D1, D2..., Ds), each containing 48/S ITSCF conditions. Within each shard, the data are further divided into R slices such that every slice includes all six fault labels, ensuring label balance and diversity. To better reflect practical deployment, fault cases with similar severity levels were grouped into the same shard. When a contaminated or poisoned data point is found in a specific condition within one of these shards (e.g., D1), the corresponding data shard D1 is regarded as poisoned. Upon detection of such corrupted data, only the affected shard is retrained from scratch, while the remaining shards and their learned parameters are preserved.

Experimental results demonstrate that SISA unlearning restores diagnostic accuracy while reducing retraining time markedly. For instance, when comparing non-SISA full retraining with poisoned data to SISA unlearning with two shards, the SISA model with two shards achieves slightly lower accuracy, from 95.69% to 99.05% when eliminating poisoned data. Furthermore, regarding computational efficiency, "when the number of shards is set to two, the SISA unlearning process reduces the retraining time to 221.8 s, and additionally decreases to 112.2 s when four shards are applied, achieving speed-ups of 2.01× and 3.97×, respectively. Overall, the proposed SISA-based MU framework can effectively eliminate the negative influence of poisoned data caused by sensor failures, while achieving comparable accuracy to full retraining at considerably lower computational cost. The results also reveal that excessive sharding may reduce accuracy due to limited data diversity, and that misclassifications mainly occur among the same phase on each side and three phases on the LV side."

In conclusion, "this paper presented a SISA-based MU framework for the power transformer ITSCF localization task considering sensor failures. By partitioning the whole dataset into multiple shards and retraining only the affected shard, the proposed method efficiently removes the impact of poisoned data while avoiding full retraining. The framework provides a practical, computationally efficient unlearning strategy for power transformer ITSCF localization.

Improvements for AI systems

As a fastidious researcher, I have analyzed the proposed SISA-based Machine Unlearning (MU) framework for Power Transformer Inter-Turn Short-Circuit Fault Localization (ITSCF). The core strength of this work lies in its ability to perform selective data removal without full retraining, offering a critical efficiency gain in real-world industrial monitoring.

Here are the specific improvements and capabilities this research enables for AI systems:


)

  1. Enhanced Robustness Against Data Poisoning (Sensor Failures):

The primary improvement is the integration of the SISA architecture directly into the training/inference pipeline to enable targeted unlearning. The system can now proactively identify and remove specific poisoned data points (e.g., those caused by EMI sensor failures) from a pre-trained model without discarding all prior knowledge.

  1. Significant Reduction in Computational Overhead for Model Updates:

The system drastically reduces the computational cost associated with model maintenance. Instead of performing full retraining (which is computationally expensive and time-consuming), the framework allows for targeted retraining only on the specific poisoned shards identified by a fault detection mechanism. This translates to orders of magnitude faster update cycles in dynamic industrial environments.

  1. Adaptive Model Maintenance Based on Fault Localization:

The improved AI system can now perform just-in-time model refinement. When a sensor failure is detected (e.g., via external diagnostics), the framework immediately isolates the relevant data shard and retrains only that portion, ensuring the model remains accurate for current operating conditions while efficiently purging the influence of faulty sensor readings.

  1. Preservation of General Knowledge During Unlearning:

Unlike simple deletion methods, this system preserves the knowledge learned from healthy operational data contained in unaffected shards. This prevents catastrophic forgetting—where removing one type of contaminated data causes the model to forget how to classify other, valid fault conditions (e.g., HA vs. HB).

  1. Improved Diagnostic Accuracy Under Contaminated Conditions:

Experimental results show that SISA unlearning restores diagnostic accuracy close to that of full retraining (e.g., achieving >97% accuracy after removing poisoned data). This means the deployed AI system maintains high reliability even when faced with realistic, noisy sensor data, leading to more trustworthy fault localization decisions in critical infrastructure.

  1. Optimized Model Scaling via Sharding Strategy:

The research provides empirical insight into how the number of shards influences performance. The system can now dynamically select the optimal sharding strategy (e.g., S=2 vs. S=4) based on the available data complexity and required computational budget, balancing accuracy against speed gains, a crucial feature for large-scale deployment planning.

Related papers