Procedural Knowledge Extraction from Industrial Troubleshooting Guides Using Vision Language Models

arXiv:2601.22754 · cs.CV, cs.AI · Submitted 2026-08-23 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Procedural Knowledge Extraction from Industrial Troubleshooting Guides Using Vision Language Models".

Jane: The paper was written by the authors from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary: Tom: So, the paper starts by summarizing the problem and then outlining their approach in "Procedural Knowledge Extraction from Industrial Troubleshooting Guides Using Vision Language Models." They establish that these guides are rich in expertise but are fundamentally challenging to digitize because of how they rely on visual flow.

Jane: The summary makes it clear that traditional methods fail because the procedural knowledge—the steps and the branches—is conveyed through symbols like diamond shapes and arrows, not just through linear sentences. It’s a visual language that is difficult for simple text pars to understand.

Lu: I agree with Jane; this is a classic case of linguistic versus spatial reasoning. The paper highlights that current AI solutions are usually trained on sequential discourse, like recipes or step-by-step textual instructions, which doesn't apply here where the logic dictates the flow.

Meng: And Meng wants to emphasize that this isn's just a minor technical hurdle for automation; it’s a significant challenge because we have millions of these guides across different manufacturers and trying to automate them is impossible with current methods without addressing this visual intelligence gap.

Lalam: Lalam sees that the summary shows us exactly where the AI needs to operate. We need the AI to understand the entire flow of process, not just read isolated text fragments, which is critical for moving from human-level diagnostic understanding to reliable machine-level automation.

Improvements/Methodology: Tom: Moving forward in "Procedural Knowledge Extraction from Industrial Troubleshooting Guides Using Vision Language Models," we look at the methodology—specifically how the researchers designed their experiments to tackle this complex visual task. They introduced a specific schema and two distinct prompting strategies.

Jane: The schema is a very elegant solution, defining a uniform way of categorizing everything into conditions, actions, and decisions so that the models can understand the structure of the troubleshooting process without ambiguity. It helps standardize messy information.

Lu: I find their dual prompting strategy incredibly insightful; they're testing whether adding explicit visual cues—telling the model exactly what a diamond shape means functionally—can significantly improve its ability to grasp graph logic and connectivity.

Meng: In terms of execution, this is where we see the input pipeline being designed. We’re not just feeding raw images to the AI; we’re providing structured instructions about how those visual components are supposed to relate to each other, which makes a huge difference in training data quality.

Lalam: Lalam views this methodology as teaching the AI how to think spatially and connect the dots in a way that can fundamentally improve our ability to document and share industrial processes. It moves us toward building a robust digital twin of knowledge itself, allowing us to improve our organizational culture through shared clarity.

Results: Tom: Now, looking at the results in "Procedural Knowledge Extraction from Industrial Troubleshooting Guides Using Vision Language Models," we see the quantitative performance of Qwen2-VL-7B and Pixtral-12B under those two prompting strategies. The findings are pretty sobering, showing limited overall capability.

Jane: It's important to teach our listeners that while Qwen achieved higher peak performance under the standard approach, the most critical part—the relation extraction—was terrible for both models, scoring below zero point one one F1. That poor connection rate is a massive indicator of failure in understanding the flow.

Lu: This disparity in results is really telling; it suggests that simply having a larger model like Pixtral doesn't automatically solve the problem if its underlying architecture isn't designed to track those dense visual connections. The size doesn's not always enough when you’re dealing with spatial logic.

Meng: An engineer would point out that we can't just throw a bigger model at this problem; the performance gap is rooted in how these architectures process the visual information, not just in their parameter count. It' structural design is what makes the a difference.

Lalam: And Lalam sees these results as an incredibly honest look at where our current AI capabilities are failing us—it’s a very clear picture of the limitations we face when trying to automate complex reasoning tasks.

Discussion/Analysis: Tom: The discussion section of "Procedural Knowledge Extraction from Industrial Troubleshooting Guides Using Vision Language Models" digs into the specific failure modes, which is where things get truly interesting regarding why these models struggled with the visual data.

Jane: The authors describe a unique issue in Qwen2-VL called "infinite loop collapse," where it starts generating hundreds of near-identical entities, which is a major problem for reliable data extraction because it’s not capturing the true process.

Lu: And Pixtral shows its own pathology—what they call capacity saturation—where it simply lacks the necessary capacity to track all the interconnected parts of a complex flowchart, even though it’s larger than Qwen. The complexity overwhelms its internal logic.

Meng: A practical concern that both models failed at is parsing the arrows and connections between nodes. They found individual components, but they couldn't read the flow, which is critical for building a proper knowledge graph of operations.

Lalam: Lalam finds this failure mode analysis incredibly valuable because it tells us precisely where to focus our development efforts. We can't just assume that knowing *what* to look for is enough; we need to understand *how* the AI thinks and how its internal mechanism works.

Conclusion/Wrap-up: Tom: So, as we conclude this entire conversation about "Procedural Knowledge Extraction from Industrial Troubleshooting Guides Using Vision Language Models," the authors confirm that while these VLMs are promising, they are currently far from being ready for autonomous industrial deployment.

Jane: They emphasize again that the relation extraction bottleneck is a fundamental issue not solvable with prompting alone, which is a sobering realization for any developer building these systems into reality.

Lu: We must acknowledge the limitations—like the 40GB VRAM constraint and single-manufacturer data—but also see the path forward through targeted improvements like constrained decoding to address those specific failures.

Meng: I believe the real-world implication here is that this technology is perfectly suited for human-AI collaboration, allowing AI to pre-populate data fields while experts verify failures, rather than replacing human expertise entirely.

Lalam: Lalam sees this entire journey as a necessary milestone in the development of our industrial knowledge base. We’ve quantified the gap and provided a clear path toward how AI can better support our shared industrial expertise in a culture of continuous improvement.

cs.CV, cs.AI

Submitted: 2026-08-23

Updated: 2026-08-25

Importance score: 82/100

The gist: The paper, "Procedural Knowledge Extraction from Industrial Troubleshooting Guides Using Vision Language Models," investigates the feasibility of automating the extraction of structured procedural

Key concepts

Procedural Knowledge Extraction
This is the core task of pulling specific steps and branching logic from complex industrial guides. It involves understanding how a machine or process moves through various decisions, rather than just reading a simple list of instructions.
Relation Extraction Failure
This refers to the specific technical failure where models could identify individual components of a flowchart but failed critically at understanding the connections or flow between those steps. This inability to track the sequence is a major limitation.
Dual Prompting Strategy
The researchers implemented this strategy to improve AI performance. They provided structured instructions and explicit visual cues, essentially teaching the models how to interpret specific shapes and their functional relationships within the troubleshooting process.

Terminology

Summary

The paper, Procedural Knowledge Extraction from Industrial Troubleshooting Guides Using Vision Language Models, investigates the feasibility of automating the extraction of structured procedural knowledge (PK) from industrial troubleshooting guides, which rely heavily on visual elements like flowcharts and symbolic notation.

Problem and Motivation:

Industrial maintenance is a knowledge-intensive activity where guides encode domain expertise through branching decision points, symbolic notations, and dense technical language. However, these documents are challenging to process automatically due to their heterogeneous layouts, inconsistent terminology, and reliance on visual relationships. The authors aim to address this bottleneck by leveraging Vision Language Models (VLMs) to automate the extraction and structuring of this knowledge for operator support systems.

Methodology:

The study utilizes a dataset consisting of 12 proprietary industrial troubleshooting guides (24 pages total), which follow a flowchart-like style where Rectangular boxes describe observations or actions, diamond shapes represent decision points, and arrows define the procedural flow. The core task is framed as a structured output generation problem using a defined schema:

  • Entity Types: Condition (a state to be verified), Action (an operation to be performed), and Decision (a branching point with yes/no outcomes).

  • Relation Type: isPrecededBy (one step is preceded by another in the procedure).

The researchers evaluated two open-weight VLMs: Qwen2-VL-7B and Pixtral-12b. Two prompting strategies were compared:

  1. Standard Instruction-Guided Prompt: Providing only initial instructions and the schema outline.

  2. Augmented Prompt: Including explicit descriptions of the visual conventions used in the diagrams, including the functional role of shapes and the interpretation of arrows and branch labels.

The evaluation metrics included Entity Precision, Entity Recall, Relation Precision, Relation Recall, and F1 scores for both entities and relations.

Results:

The overall performance across all 12 guides was limited. Both models demonstrated low extraction capabilities:

  • Entity F1 scores ranged from 0.24 to 0.34.

  • Relation F1 scores were consistently below 0.11.

Qwen2-VL achieved the highest entity performance (F1 = 0.340 under standard prompting). However, relation extraction was a severe bottleneck for both models, with Qwen2-VL extracting only 5% recall of ground truth relations and Pixtral performing worse.

Prompt engineering yielded divergent effects:

  • Qwen2-VL: Augmented prompting improved relation extraction (F1: 0.061 to 0.107) but reduced entity precision (0.305 to 0.203).

  • Pixtral-12B: Showed degradation under augmented prompting for both entities and relations, indicating that increased prompt complexity may conflict with its processing architecture.

Discussion and Failure Modes:

The limitations stemmed from distinct failure modes:

  1. Qwen2-VL exhibited infinite loop collapse, where documents failed completely (F1 0.15) by entering repetitive loops generating hundreds of near-identical entities with incrementing identifiers.

  2. Pixtral exhibited capacity saturation, achieving only 2–7 valid relations total despite having a larger parameter count than Qwen, suggesting it struggle[s] when tracking the overlapping arrows and tightly packed blocks of Dutch text that characterize these diagrams.

A critical shared failure was the relation extraction bottleneck. Both models often identified nodes but failed to parse the connecting arrows. This suggests that parsing spatial relationships in diagrams with overlapping arrows and densely packed nodes exceeds current architectural or training capabilities.

Conclusion:

The study concludes that current open-source VLMs are unsuitable for autonomous deployment in safety-critical settings due to high entity and relation miss rates, unpredictable failure modes, and the inability to reliably parse spatial relationships. While the results are far short of requirements for autonomous knowledge graph construction, documents without hallucinations reached entity F1 scores between 0.48 and 0.78, suggesting that models could support human-in-the-loop workflows to reduce annotation time.

Improvements for AI systems

Based on a rigorous analysis of the identified failure modes—specifically relation extraction bottlenecks, infinite loop collapse, and poor spatial reasoning—the following engineering improvements are necessary to elevate current VLM performance for industrial procedural knowledge extraction.


We must move beyond simple sequence-to-JSON output and explicitly mandate the structural representation of the flowchart.

  • Technical Improvement: Integrate a Graph Neural Network (GNN) layer into the VLM's decoding pipeline. Instead of merely listing entities, the model must be trained to identify nodes (entities) and then predict directed edges (relations).

  • Specific Mechanism: The output should not be a flat list of (E 1, E 2) pairs, but a structured representation: Graph = Nodes: [E 1, E 2,...], Edges: [(E i, E j),.... This forces the model to prioritize connectivity over mere entity enumeration.

The failure of models like Pixtral to track overlapping arrows and densely packed blocks indicates a deficiency in spatial coherence, not just visual recognition.

  • Technical Improvement: Implement Spatial-Flow Attention Mechanisms. The VLM must be trained to compute attention weights that are dependent on the geometric relationship between two detected entities (e.g., proximity and directional alignment) rather than just semantic overlap.

  • Specific Mechanism: When processing a flowchart segment, the model must maintain a localized flow state vector, dynamically adjusting its resolution (similar to Qwen2-VL's approach but more robust) to ensure that the path of an arrow—even when it crosses multiple nodes—is maintained as a single, continuous relational path.

The infinite loop collapse observed in Qwen2-VL is a critical safety flaw for autonomous systems.

  • Technical Improvement: Introduce Sequence and State Machine Validators post-inference. The output must be validated against the expected constraints of the procedural schema (e.g., a state cannot transition to itself, or an action cannot precede a condition).

  • Specific Mechanism: Implement a Repetition Monitor that detects sequences where identical entity text appears multiple times with incrementing identifiers (e.g., E 16, E 17,... all reading Zijn er verkeerde doppen aanwezig?). If repetition exceeds a threshold, the generator must terminate and flag that specific output as invalid, preventing token budget exhaustion.

The initial instruction-guided prompt is insufficient to overcome architectural limitations; visual guidance must be formalized.

  • Technical Improvement: Develop Visual Schema Prompts (VSPs). The system will receive a structured preamble describing the functional role of every element in the guide (e.g., Diamond shapes represent conditional branching, Rectangles are terminal actions).

  • Specific Mechanism: This VSP integrates domain-specific visual conventions directly into the input token stream, allowing the model to prioritize structural interpretation before semantic content extraction.

The resulting system will transcend simple data extraction and become a Procedural Knowledge Graph Generator (PKGG) capable of:

  1. Autonomous Fault Diagnosis: By accurately constructing the required flow graph, the system can take sensor inputs (matching Conditions) and trace the precise path through branching Decisions to recommend only verified Action sequences, replacing human guesswork.

  2. Dynamic Procedural Simulation: The PKGG can simulate the troubleshooting process in real-time, highlighting which branch must be taken based on current sensor data and providing a complete visualization of the required sequence of operations before the operator acts.

  3. Error Detection and Validation: The system will automatically flag any document where the extracted graph violates logical flow constraints (e.g., an action appearing before a necessary decision point), ensuring that maintenance procedures are logically sound, even if they originate from heterogeneous or poorly documented sources.

  4. Augmented Human-AI Collaboration: By providing high-confidence structural outputs (the GNN structure) and flagging low-confidence or failed extractions (via the loop/validation monitors), it allows human experts to rapidly validate machine output, significantly reducing annotation time and accelerating deployment in critical industrial settings.

Sources

Related papers