Knowledge-Graph-Guided Retrieval-Augmented LLMs for Explainable Root Cause Analysis in Automotive HiL Validation
Hamza Ouarrad, Mohammad Abboush, Andreas Rausch
Technische Universität Clausthal
cs.CR, cs.SE
Submitted: 2026-08-11
Updated: 2026-08-13
Comments: 10 pages, 3 figures, 5 tables. Accepted for publication and oral presentation at the 10th International Conference on System Reliability and Safety (ICSRS 2026), Rome, Italy, November 23--25, 2026
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: This paper proposes a knowledge-graph-guided retrieval-augmented large language model (KG-guided RAG-LLM) framework for root cause analysis (RCA) and fault localization in automotive
Terminology
Summary
This paper proposes a knowledge-graph-guided retrieval-augmented large language model (KG-guided RAG-LLM) framework for root cause analysis (RCA) and fault localization in automotive Hardware-in-the-Loop (HiL) validation data. The method addresses the limitation that "Industry-certified tools predominantly rely on rule-based analysis and predefined thresholds, which are effective for flagging anomalies but provide little support for identifying the actual location and root cause of a fault. The authors note that
an anomaly observed in a single signal may originate from sensors, communication channels, software components, controllers, or actuators [5], and the corresponding fault effects often propagate through interconnected components before becoming visible in the recordings [6]."
The framework operates as follows: raw multivariate HiL recordings are first converted into compact diagnostic evidence, including abnormality scores, signal reaction times, deviation directions, and inter-signal correlation changes.
This evidence is then enriched with sensor-to-location and propagation knowledge
from a knowledge graph, and similar historical cases are retrieved to support the final reasoning step. The LLM is used as a decision and explanation layer rather than as a direct time-series classifier, producing a ranked fault-location prediction together with an interpretable RCA explanation.
The authors claim this is "the first study that unifies automotive HIL recordings, structured diagnostic evidence, case-based retrieval, knowledge-guided diagnosis, root cause analysis, and Retrieval-Augmented Large Language Models within a single framework."
The methodology involves several key stages. First, raw recordings are preprocessed and compared against time-aligned healthy reference runs, with the analysis focusing on a pre-fault interval
(2–11 s) as baseline and a stable fault interval
(15–45 s) for abnormal response. For each signal, the framework computes raw mean difference, mean-shift z-score, maximum-deviation z-score, first abnormal reaction time, and an overall abnormality score.
A knowledge graph links each candidate fault location to its physically related signals,
containing signal-to-fault mappings, direct and propagation relationships, and diagnostic rules.
A KG-guided evidence score is computed as Sevid(c) = wp Sp(c) + wd Sd(c) + wg Sg(c), where the weights give the strongest importance to primary and direct KG evidence, while propagation signals provide supporting evidence only.
Retrieval uses cosine similarity on structured z-score vectors, with a leave-one-recording-out strategy
to prevent retrieval leakage. The final candidate ranking combines KG-guided evidence with retrieval support via S(c) = αSevid(c) + (1−α)Sret(c). The LLM receives a compact diagnostic report rather than raw time-series input
and is instructed to return valid JSON containing the predicted Top-1 fault location, the Top-3 ranked candidate list, a confidence level, candidate-level reasoning, the most likely propagation path, uncertainty notes, and a short RCA explanation.
Window-level predictions are aggregated into recording-level diagnoses using rank-based voting over the Top-3 lists.
The framework was evaluated on two HiL case studies: an ASM gasoline engine system and an electric vehicle (EV) system, both using dSPACE SCALEXIO real-time platforms with MicroAutoBox II controllers. The ASM case study considers five fault locations (accelerator pedal, brake pedal, steering wheel, engine speed, throttle position) with two faulty recordings per location (ten total), while the EV case study considers three fault locations (high-voltage battery, rear electric machine speed, steering system) with one recording each (three total). The fault type is a gain fault, active approximately between 12 s and 47 s.
Results show that Gemma-3 27B and Qwen3 32B achieve the strongest performance
on both case studies. On the ASM gasoline engine case, these models achieve a Top-1 accuracy of 0.900, a Top-3 accuracy of 1.000, and an MRR of 0.933.
On the EV case, they reach a Top-1 accuracy of 0.944
with Top-3 accuracy of 1.000. Recording-level aggregation further improves results: "In the ASM gasoline engine case, the window-level Top-1 accuracy is 0.833, while the file-level Top-1 accuracy reaches 1.000. Similarly, in the EV case, the window-level Top-1 accuracy is 0.944, and the file-level Top-1 accuracy also reaches 1.000. The retrieval component was independently evaluated, with
the structured z-only diagnostic representation achiev[ing] the strongest performance on the ASM gasoline engine case study, with 0.50 Top-1 accuracy and 1.00 Top-3 accuracy."
Regarding computational complexity, the authors report that Gemma-3 27B requires 4.773 s per prompt, whereas Qwen3 32B requires 6.614 s per prompt,
meaning Gemma reaches the same highest localization performance while reducing the average inference time by approximately 28%.
The authors conclude that Gemma-3 27B provides the most favorable accuracy-latency trade-off and is selected as the preferred model for the proposed KG-guided RAG-based fault-localization pipeline.
The paper acknowledges limitations: "The evaluation dataset is still limited in size and mainly focuses on gain faults. In addition, the present experiments address single-fault localization, while concurrent faults may create overlapping direct and propagated effects that are more difficult to separate. The quality of the RCA explanation also depends on the completeness of the knowledge graph and the available historical case library. Future work will
extend the framework to concurrent and interacting fault scenarios, larger HiL datasets, and additional fault types, as well as
integrate richer spatial or causal representations, such as dynamic knowledge graphs or graph neural networks, to model fault propagation more explicitly."
Improvements for AI systems
Improvements to AI Systems Based on This Paper
-
Hybrid Diagnostic Reasoning Layer: Integrate a KG-guided evidence scoring mechanism (combining primary, direct, and propagation signal weights) into the LLM's reasoning loop, enabling the AI to prioritize physically grounded signal relationships over statistical correlations alone. The improved system can localize faults in complex cyber-physical systems by explicitly modeling fault propagation paths, rather than relying solely on pattern matching.
-
Structured Evidence Preprocessing for Time-Series: Replace raw time-series input with compact diagnostic evidence vectors (abnormality scores, reaction times, deviation directions, correlation changes) before LLM inference. This reduces token complexity and improves reasoning accuracy by presenting only causally relevant features. The system can now handle multivariate sensor data from industrial validation rigs without overwhelming the model with noise.
-
Case-Based Retrieval Augmentation with Leakage Prevention: Implement a retrieval module using cosine similarity on structured z-score vectors with a leave-one-recording-out strategy to avoid data leakage. The enhanced system can leverage historical fault cases to improve diagnosis accuracy, achieving 100% Top-1 accuracy at file-level aggregation, while maintaining generalizability to unseen recordings.
-
Rank-Based Temporal Aggregation: Use window-level predictions aggregated via rank-based voting over Top-3 candidate lists to produce recording-level diagnoses. This enables the AI to combine evidence across time windows, improving robustness against transient anomalies and yielding higher confidence in final fault localization (e.g., from 0.833 to 1.000 Top-1 accuracy in the ASM case).
-
Interpretable JSON Output with Uncertainty Quantification: Instruct the LLM to return structured JSON containing ranked candidates, confidence levels, propagation paths, and uncertainty notes. This makes the AI's decision-making auditable and actionable for engineers, allowing them to verify reasoning steps and trust the system in safety-critical automotive validation.
-
Knowledge-Graph-Guided Signal-to-Location Mapping: Enrich the LLM's context with a knowledge graph that links each fault location to physically related signals and propagation rules. The improved system can now reason about indirect fault effects (e.g., a sensor anomaly caused by a downstream actuator failure) and provide more accurate root cause explanations than rule-based tools.
-
Model Selection for Latency-Accuracy Trade-off: Adopt a two-model strategy where a lighter model (e.g., Gemma-3 27B) is preferred when inference speed matters, given its 28% faster inference with identical accuracy compared to larger alternatives. The system can dynamically choose between models based on real-time constraints, making it deployable in HiL environments with strict timing requirements.
-
Fault Propagation Path Prediction: Enable the LLM to output the most likely propagation path (e.g., sensor → controller → actuator) by combining KG edges with diagnostic evidence. This allows the AI to not only locate the fault but also explain how it spread through the system, aiding in rapid root cause mitigation and future design improvements.
-
Single-Fault to Multi-Fault Extension Readiness: Structure the evidence scoring and retrieval to be modular, allowing future extension to concurrent fault scenarios by separating overlapping direct and propagated effects. The improved system can be adapted to diagnose multiple simultaneous faults once more training data is available, without redesigning the core pipeline.
-
Cross-Domain Transferability: The framework's separation of raw data preprocessing, KG-based evidence enrichment, and LLM reasoning makes it applicable to other industrial validation domains (e.g., aerospace, energy systems) where multivariate sensor data and known component relationships exist. The improved AI can be retrained with minimal effort for new domains by updating the knowledge graph and historical case library.
Abstract
Hardware-in-the-Loop validation of automotive software systems generates large multivariate time-series recordings whose manual analysis is time-consuming and often limited to anomaly detection and fault classification rather than root-cause analysis. Although deep learning methods have shown strong performance in fault detection and classification, they usually require task-specific training or retraining when new fault locations, systems, or operating conditions are introduced. They also tend to treat localization as a classification task, without explicitly representing the spatial and functional relationships between fault locations, sensors, and downstream subsystem effects. This limits their generalizability and their usefulness for engineering root cause analysis and diagnosis. This paper proposes a knowledge-graph-guided retrieval-augmented large language model framework for RCA (root cause analysis) and fault localization in automotive HiL data. The method converts raw time-series recordings into compact diagnostic evidence, enriches this evidence with sensor-to-location and propagation knowledge, and retrieves similar historical cases to support the final reasoning step. The LLM is then used as a decision and explanation layer rather than as a direct time-series classifier, producing a ranked fault-location prediction together with an interpretable RCA explanation. The framework is evaluated on two automotive HiL case studies: an ASM gasoline engine and an electric vehicle system. The best-performing model achieves Top-1 accuracies of 90% and 94%, respectively, while recording-level aggregation reaches perfect file-level fault localization in the evaluated subset. These results demonstrate the potential of KG-guided RAG-LLM reasoning for explainable and generalizable HiL RCA.
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs