Comparing Object Detection Models for Electrical Substation Component Mapping
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Comparing Object Detection Models for Electrical Substation Component Mapping".
Jane: The gist The research trains and compares three object detection models (YOLOv8, YOLOv11, RF-DETR) to map key electrical substation components in US images,
Tom: First, who's behind it and why it matters.
Paper summary: Jane: So, wrapping up the discussion on "Comparing Object Detection Models for Electrical Substation Component Mapping," what's the real takeaway for listeners who might be interested in this kind of technology?
Tom: The main thing to remember is that using AI models like YOLOv8 can be the best starting point for automated substation mapping based on this study because it showed the highest average mean average precision.
Lu: And we have to remember the specific comparison they made: YOLOv11 was the fastest in training time, but it didn't translate into better accuracy than YOLOv8.
Meng: So, for an engineer looking at this paper, what’s the practical application right now? Is this ready to implement on a substation next week?
Tom: No, it’s not ready for deployment yet because the authors themselves admit that even with their work, they still have a relatively small dataset.
Jane: So what's the final thought on how this research fits into the bigger picture of using AI for critical infrastructure?
Lalam: This paper demonstrates that object detection models can definitely begin to recognize substation components when you’re working under constraints like a limited dataset, and that's a step in the right direction.
Conclusion: Tom: So we're wrapping up this look at "Comparing Object Detection Models for Electrical Substation Component Mapping." This paper takes three different AI models, YOLOv8, YOLOv11, and RF-DETR and sees how well they can find parts in substation pictures.
Jane: Exactly. The authors are trying to figure out which of these tools is actually the most reliable way to automatically map those complex electrical grids.
Lu: It’s interesting because they look at standard deep learning techniques, like CNNs, and apply them directly to something super practical and dangerous—power infrastructure.
Meng: Practically speaking, they’re moving away from needing a person on site for every single inspection. That saves a ton of time and reduces the risk to workers out there.
Lalam: This research shows that even with limited data, these models can start recognizing specific components like transformers and circuit breakers in real-world images.
Tom: The big result they found is that YOLOv8 ended up being the top performer among the three models they tested for finding reactors and circuit breakers.
Jane: And while YOLOv11 was quicker to train, it didn't beat YOLOv8 in terms of how accurate it actually was at detecting those specific items.
Lu: The authors mention that the performance difference between the models isn't huge on smaller components, but it gets pretty significant when you look at the overall map they generate.
Meng: So, for an engineer listening to this, what does that mean right now? Is this ready to go onto a real grid monitoring system?
Tom: Not yet. They are being very clear that because the dataset is relatively small, these results aren't quite ready for real-world deployment on live equipment.
Jane: But the paper's point isn't that it’s perfect today; it’s showing us a viable path forward for automated mapping when we don't have perfect, massive datasets available.
Lalam: It shifts the focus toward how we can use this kind of AI to build better safety and maintenance tools, even when the initial data isn't ideal.
Tom: Because what they show is that these object detection models are already capable of doing a lot more than just identifying things—they can start building that map.
Jane: And next time we talk about this, we’ll be looking at how to make those datasets bigger and better so this tech can actually move from the lab to the power lines.
cs.CV
Submitted: 2025-12-27
Updated: 2026-10-08
Comments: 42 pages, 6 figures
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 71/100
The gist: The gist The research trains and compares three object detection models (YOLOv8, YOLOv11, RF-DETR) to map key electrical substation components in US images, aiming to find the most reliable method
Key concepts
- Object Detection Models
- These are deep learning algorithms trained to identify and locate specific objects within an image. They work by analyzing pixels to draw bounding boxes around items of interest, such as transformers or circuit breakers, allowing computers to automatically find these components in photographs.
- YOLOv8
- This is a specific version of the You Only Look Once (YOLO) algorithm used in the study. It is known for being very fast and efficient during inference, meaning it can process images quickly without sacrificing much accuracy, making it a strong candidate for real-time applications.
- mAP (mean Average Precision)
- This is a standard metric used to evaluate how well an object detection model performs. It measures the accuracy of the model's predictions across all classes. A higher mAP score indicates that the model is more precise and reliable in identifying substation components.
- Data Augmentation
- This technique involves artificially increasing the size and diversity of a training dataset by applying various transformations to existing images. In this study, they used rotations and hue adjustments to create 875 training images per component, helping the models generalize better to different lighting or orientations.
Terminology
Summary
The gist The research trains and compares three object detection models (YOLOv8, YOLOv11, RF-DETR) to map key electrical substation components in US images, aiming to find the most reliable method for automated mapping.
Introduction and Problem Statement
Electrical substations are pivotal nodes in the electrical grid that step up or step down voltage levels, and their failure can have significant economic and public safety implications The importance of substations and their components goes beyond standard maintenance and regulation because any failure can have significant economic and public safety impacts Traditional manual mapping of substation components is time-consuming and labor-intensive, requiring skilled personnel to inspect each site in person. The study investigates whether state-of-the-art computer vision models can be trained to accurately identify and differentiate between key substation components from high-resolution imagery.
Literature Review of Object Detection Models
The literature review compares how various research papers approach object detection for power infrastructure modeling. Convolutional Neural Networks, or CNN algorithms, are deep learning algorithms used for object detection, segmentation, and image classification. CNNs consist of five layers: a convolution layer, a pooling layer, activation functions, a fully connected layer, and loss functions. Overall CNNs have demonstrated strong performance in detecting substation infrastructure. Many research studies utilize the You Only Look Once (YOLO) family of algorithms utilizing a CNN backbone. YOLO models are often utilized for substation infrastructure detection due to higher efficiency, such as YOLOv8 having a higher inference speed without significant cuts in accuracy and precision. Advanced examples include fusing the YOLOv8 algorithm with an enhanced small-object detection head to improve accuracy in unexpected environments. A newly released model, Roboflow Detection Transformer (RF-DETR), is another focus of this research paper due to its limited number of studies evaluating its performance.
Model Training and Data Preparation
The methodology involves several key steps for model training, including data collection, pre-processing, and augmentation. The dataset was built using open-source imagery of U.S. substations and transmission lines. A dataset of 750 images was collected, with 250 for each component (transformers, circuit breakers, reactors). Preprocessing involved resizing and auto-orienting the images to a fixed size of 640x640 using letterbox padding. Data augmentation included rotations of 15°, 30°, -15°, and-30° degrees, resulting in 875 training images per component. Furthermore, the dataset was expanded by randomly adjusting the hue of the images to create variations in color tint up to 15 degrees.
Model Comparison and Results
The study trained YOLOv8, YOLOv11, and RF-DETR on the prepared dataset. The results showed that machine learning models have a significantly lower average accuracy detecting reactors, over circuit breakers and transformers. Of the three components attempted to detect, circuit breakers were the most accurately detected in each model with the highest average mAP. YOLOv8 achieved the highest average mAP among the three models tested, reaching an average of 0.610. The efficiency comparison showed that YOLOv11 was the most efficient model in terms of training time, but this did not translate into higher average accuracy.
Conclusion and Implications
The research concludes that automated substation component mapping is more effective than a manual approach because it can process many images quickly and save much time. Although the resulting model accuracies would not yet support real-world deployment due to the relatively small dataset, the work demonstrates that object detection models can still begin to recognize substation components under these constraints. The YOLOv8 model is identified as the best model for this scenario due to its higher average mAP, indicating it performed best in terms of overall detection accuracy. This research serves an important purpose in understanding the benefits of utilizing automated substation mapping rather than manual. By continuing to explore these AI-driven approaches, we can enable more informed responses to dangers. The final results showed a count of 7615 circuit breakers, 3132 transformers, and 1133 reactors labeled through the inference mapping done on these images.
Bibliography
Alemazkoor, N., Ayyalasomayajula, A., & Li, M. (2021). Mapping electrical substations using deep learning and overhead imagery (arXiv:2104.06601). arXiv. https://doi.org/10.48550/arXiv.2104.06601
Alzubaidi, L., Zhang, J., Humaidi, A. J., Duan, Y., Santamaría, J., Fadhel, M. A., & Farhan, L. (2021). Review of deep learning: Concepts, CNN architectures, challenges, applications, future directions. Journal of Big Data, 8(53). https://doi.org/10.1186/s40537-021-00444-8
Chai, S., & Lau, V. K. N. (2021). Multi-UAV trajectory and power optimization for cached UAV wireless networks with energy and content recharging: Demand-driven deep learning approach. IEEE Journal on Selected Areas in Communications, 39(10), 3208–3224. https://doi.org/10.1109/JSAC.2021.3088694
Erenoğlu, A. K., Sengor, I., & Erdinç, O. (2024). Power system resiliency: A comprehensive overview from implementation aspects and innovative concepts. Energy Nexus, 15(100311). https://doi.org/10.1016/j.nexus.2024.100311
Faisal, M. A. A., Mecheter, I., Qiblawey, Y., Fernandez, J. H., Chowdhury, M. E., & Kiranyaz, S. (2025). Deep learning in automated power line inspection: A review. Applied Energy, 385(125507). https://doi.org/10.1016/j.apenergy.2025.125507
Florkowski, M., Kuniewski, M., Furgal, J., & Pajak, P. (2017). Investigation of overvoltages in distribution transformers. In Proceedings of the 18th International Scientific Conference on Electric Power Engineering (EPE) (pp. 1–4). https://doi.org/10.1109/EPE.2017.7967288
Gallagher, J. (2024). A multispectral automated transfer technique (MATT) for machine-driven image labeling utilizing the Segment Anything Model (SAM) [Data set]. IEEE DataPort. https://doi.org/10.21227/g06c-yh08
Gallagher, J. E., & Oughton, E. J. (2023). Assessing thermal imagery integration into object detection methods on air-based collection platforms. Scientific Reports, 13(8491). https://doi.org/10.1038/s41598-023-34791-8
Gallagher, J. E., & Oughton, E. J. (2024). VORTEX: A spatial computing framework for optimized drone telemetry extraction from first-person view flight data (arXiv:2412.18505). arXiv. https://arxiv.org/abs/2412.18505
Gallagher, J. E., & Oughton, E. J. (2025). Surveying You Only Look Once (YOLO) multispectral object detection advancements, applications, and challenges. IEEE Access, 13(8654–8684). https://doi.org/10.1109/ACCESS.2025.3526458
Girshick, R. (2015). Fast R-CNN (arXiv:1504.08083). arXiv. https://doi.org/10.48550/arXiv.1504.08083
Girshick, R., Donahue, J., Darrell, T., & Malik, J. (2013). Rich feature hierarchies for accurate object detection and semantic segmentation (arXiv:1311.2524). arXiv. https://doi.org/10.48550/arXiv.1311.2524
Hall, J. W., Tran, M., Hickford, A., & Nicholls, R. J.
Improvements for AI systems
-
Bold model selection for component mapping: Select YOLOv8 as
the best model to use for this scenario
because it achievedthe highest mean average precision, or mAP,
indicating itperformed best in terms of overall detection accuracy.
-
Bold component-specific performance analysis: Identify that
circuit breakers were the most accurately detected in each model with the highest average mAP,
while noting thatreactors performed significantly worse in each model, reaching a maximum mAP of around 0.45
due to their appearance varying. -
Bold real-world deployment capability: The improved system can be utilized for
automated substation component mapping
by leveraging the inference mapping done on collected images to provide aconcrete number of each component type,
such asCircuit Breakers 7615.
-
Bold efficiency optimization strategy: Implement a decision rule that favors accuracy over speed, as the paper found that
efficiency is not a major concern compared to accuracy
for this use case, concluding that YOLOv8 is the optimal choice despite slower training times compared to YOLOv11. -
Bold data constraint acknowledgment: The system should incorporate a mechanism to acknowledge current limitations, stating that
Because our dataset was relatively small and contained lower-quality images, the resulting model accuracies would not yet support real-world deployment,
positioning the output as aproof of concept.
Abstract
Electrical substations are a significant component of an electrical grid. Indeed, the assets at these substations (e.g., transformers) are vulnerable to hazards such as hurricanes, flooding, earthquakes, and geomagnetically induced currents (GICs). Because failures can have significant economic and public safety implications, identifying key substation components is essential for quantifying vulnerability. Unfortunately, traditional manual mapping of substation infrastructure is time-consuming and labor-intensive. Therefore, an autonomous solution utilizing computer vision models is preferable, as it offers greater convenience and efficiency. In this study, we train and compare 16 models on a manually labeled dataset of US substation images. These models include 12 You Only Look Once (YOLO) models, 2 Roboflow Detection Transformer (RF-DETR) models, and 2 Cascade R-CNN models. RF-DETR-large achieved the highest overall detection performance with mAP@50 and mAP@50:95 scores of 0.881 and 0.632, respectively. Across all models, alternate energy systems were detected most accurately, while transformers and reactors were more difficult to identify due to their smaller size and greater visual variability. Applying our best-performing model to nationwide imagery yielded approximately 22,591 component detections across 11,083 unique substations within the United States. These detections were broken down by state and Federal Energy Regulatory Commission (FERC) regions, with Florida (2,478 detections) and Midcontinent Independent System Operator (MISO; 4,329 detections) having the largest number of detections in their respective categories.
Sources
- Zero-Shot Instance Segmentation
- VORTEX: A Spatial Computing Framework for Optimized Drone Telemetry Extraction from First-Person View Flight Data
- Fast R-CNN
- Rich feature hierarchies for accurate object detection and semantic segmentation
- Enhancing Power Grid Inspections with Machine Learning
- Major Space Weather Risks Identified via Coupled Physics-Engineering-Economic Modeling
- A Reproducible Method for Mapping Electricity Transmission Infrastructure for Space Weather Risk Assessment
- Assessing the economic benefits of space weather mitigation investment decisions: Evidence from Aotearoa New Zealand
- RF-DETR: Neural Architecture Search for Real-Time Detection Transformers
- RF-DETR Object Detection vs YOLOv12 : A Study of Transformer-based and CNN-based Architectures for Single-Class and Multi-Class Greenfruit Detection in Complex Orchard Environments Under Label Ambiguity
- Ultralytics YOLO Evolution: An Overview of YOLO27, YOLO26, YOLO11, YOLOv8, and YOLOv5 Object Detectors for Computer Vision and Pattern Recognition
- Infrared image identification method of substation equipment fault under weak supervision
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models