Constructing Dynamic Master Logic Models as Knowledge Graphs for Complex System Diagnostics Using Retrieval-Augmented Large Language Models

arXiv:2608.12304 · cs.AI · Submitted 2026-08-12 · Read on arXiv

Saman Marandi, Yu-Shu Hu, Mohammad Modarres

University of Maryland · DML Inc.

cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: 36 Pages, 8 Figures

Code: https://github.com/neo4j/neo4j

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

Terminology

Summary

Summary

This paper introduces a framework for the automated construction of Dynamic Master Logic (DML) models from system documentation, representing them as Knowledge Graphs (KG-DML) using Retrieval-Augmented Generation (RAG) and Large Language Models (LLMs). The work addresses the limitation that DML construction traditionally relies on manual, expert-driven processes, which limits scalability for complex systems. The framework is designed to transform unstructured technical documentation into executable functional models for diagnostic and reliability analysis.

The proposed approach performs model construction sequentially along the DML hierarchy, beginning with system goals and progressing through functions, subfunctions, components, and success conditions. At each layer, a retrieval query is constructed using the semantic definition of the current DML layer and the parent elements identified in the previous stage. The most relevant portions of the system documentation are retrieved and supplied to the LLM, which generates candidate elements and their logical relationships in a structured JSON format. The model is constructed using GPT-4o with temperature set to 0 for deterministic outputs. After completion of all layers, the JSON representation is converted into Cypher queries for the Neo4j graph database to create the KG-DML representation.

The framework includes three main stages: document processing and embedding, layer-by-layer model construction, and knowledge graph synthesis and export. Document preprocessing normalizes component identifiers, expands abbreviations, and resolves pronouns to reduce ambiguity. Text is segmented into chunks of 1,500 characters with 150-character overlap and embedded using text-embedding-3-small model. During construction, parent elements are processed in batches, with the impact of batch size on model quality examined in the evaluation.

The evaluation framework introduces a multi-level validation methodology combining layer-specific precision and recall, logical gate consistency analysis, and an aggregate integrity metric. The integrity score combines a weighted average of F2 scores across all layers with an exponential penalty for structural reconstruction errors, including missing nodes, missing relationships, gate mismatches, and orphan nodes. The score is defined as: IntegrityScore = 100 * (Σ w l * F l / Σ w l) * exp(-E structural / ((N nodes + N links) * S)), where F l is the F2 score for layer l, w l is the importance weight, E structural is the aggregate weighted structural penalty, and S is a penalty scaling factor.

The framework is evaluated on the Low-Pressure Coolant Injection (LPCI) safety system of a decommissioned Boiling Water Reactor, using documentation from the Interim Reliability Evaluation Program. The system includes two redundant subsystems (A and B) with multiple pumps, heat exchangers, spray headers, motor-operated valves, and injection lines. Results across five independent executions show perfect identification for Goals, Functions, and Subfunctions layers, with minor deviations at Component and Success Condition layers. For Components, F2 ranges from 0.978 to 0.990, and for Success Conditions, F2 ranges from 0.980 to 0.993.

Link-level evaluation shows perfect reconstruction for higher-level relationships (Goal-Function and Function-Subfunction), with variability concentrated at the Subfunction-Component layer where structural density is highest. At this level, recall ranges from 0.951 to 0.978, F2 ranges from 0.961 to 0.982, and gate accuracy ranges from 0.889 to 0.963. The Component-Success Condition layer exhibits near-perfect performance with gate accuracy consistently equal to 1.0.

Batch size analysis shows that a batch size of 1 achieves the highest mean integrity score of 90.38, with performance decreasing gradually as batch size increases (88.98 for batch size 5, 87.98 for batch size 10, 86.35 for batch size 15). The paper notes that a batch size of 5 provides a balance between structural consistency and computational efficiency, as reducing batch size from 5 to 1 increases LLM invocations approximately fivefold. Sensitivity analysis on the penalty scale parameter shows the integrity score responds predictably, with relative ranking across configurations remaining stable.

The resulting KG-DML supports three types of interaction: upward propagation for evaluating system success and identifying impacted upper nodes, downward propagation for generating minimal success path-sets, and explanatory queries handled through Graph-RAG. The LLM agent interprets user queries and invokes predefined graph-based tools, while probabilistic and logical reasoning remains governed by the encoded model structure. The framework supports nested Boolean gate structures, extending prior work that restricted dependencies to single flat gates.

The paper acknowledges several limitations, including evaluation on a single case study system, lack of functional equivalence testing through minimal cut sets or success path sets, no systematic ablation studies, and no explicit representation of common cause failures within the KG-DML structure. Future work includes applying the framework to multi-source documentation, modular construction across subsystems, explicit CCF modeling using beta-factor or alpha-factor approaches, functional equivalence testing, and exploring locally hosted language model deployments for confidential documentation.

Improvements for AI systems

Improvements to AI Systems:

  1. Hierarchical Knowledge Graph Construction from Unstructured Text
  • The AI system can automatically build multi-layered knowledge graphs (goals → functions → subfunctions → components → success conditions) from raw documentation, eliminating manual expert modeling.

  • It can enforce logical consistency (e.g., Boolean gate structures) across layers during generation, reducing structural errors.

  1. Retrieval-Augmented Generation with Layer-Aware Querying
  • The system can dynamically craft retrieval queries based on the semantic definition of the current modeling layer and parent elements, improving relevance of extracted information.

  • It can handle large documents by chunking (1,500 chars with overlap) and embedding, enabling scalable processing of technical manuals.

  1. Deterministic Output Generation for Engineering Applications
  • By using temperature=0 and structured JSON output, the system can produce reproducible, machine-parseable model elements—critical for safety-critical systems where variability is unacceptable.

  • It can validate outputs via layer-specific precision/recall and gate consistency checks, flagging anomalies for human review.

  1. Batch-Size Optimization for Cost-Performance Trade-offs
  • The AI can adaptively choose batch sizes (e.g., 5) to balance computational cost (LLM invocations) and structural integrity, reducing API costs by 5x compared to batch size 1 while maintaining high accuracy.
  1. Graph-Based Reasoning with Propagation and Path-Set Generation
  • The system can support upward propagation (identify impacted higher-level nodes from component failures) and downward propagation (generate minimal success path-sets) for reliability analysis.

  • It can answer explanatory queries via Graph-RAG, combining LLM interpretation with deterministic graph traversal—ensuring probabilistic reasoning stays grounded in the model.

  1. Multi-Level Integrity Scoring for Self-Assessment
  • The AI can compute a weighted integrity score (combining F2 scores and structural penalties) to self-evaluate its own model quality, enabling iterative refinement without human oversight.
  1. Normalization and Disambiguation of Technical Text
  • The system can preprocess documentation to normalize component identifiers, expand abbreviations, and resolve pronouns, reducing ambiguity that typically degrades LLM performance on engineering texts.

What the Improved AI System Can Do:

  • Automatically convert a 500-page reactor safety manual into an executable, queryable reliability model in minutes (vs. weeks of manual expert work).

  • Provide real-time diagnostic support: e.g., If pump B fails, which success conditions are still achievable? with traceable, graph-based answers.

  • Generate minimal success path-sets for complex systems with nested Boolean logic, enabling efficient fault-tree analysis.

  • Operate on confidential or proprietary documentation using locally hosted LLMs (future extension), maintaining data privacy.

  • Scale to multi-source documentation and modular subsystem construction, reducing modeling effort for large-scale industrial plants.

Abstract

Dynamic Master Logic (DML) provides a hierarchical framework for representing system behavior by linking functional objectives to underlying structural elements. However, DML construction typically relies on expert interpretation of technical documentation, limiting scalability for complex systems. This study presents a framework for automated construction of DML models from system descriptions and their representation as Knowledge Graphs (KG-DML), using Retrieval-Augmented Generation and Large Language Models as enabling tools. Building on prior work with small-scale systems, the framework extends automated KG-DML construction and evaluation to substantially larger and more complex systems. Model construction proceeds across the DML hierarchy using targeted retrieval while preserving functional dependencies and explicit logical relationships. The resulting KG-DML supports diagnostic reasoning, safety assessment, upward failure propagation, and downward dependency tracing. A multi-level validation methodology evaluates layer-specific precision and recall, logical gate consistency, and overall structural integrity. Application to the Low-Pressure Coolant Injection system of a decommissioned Boiling Water Reactor demonstrates consistent reconstruction across repeated runs. The results show that automated KG-DML construction can transform technical documentation into executable functional models for diagnostic and reliability analysis.

Sources

Related papers