GUIDE: Governed Unified Intelligence for Document-to-Artifact Generation in Enterprise Settings

arXiv:2608.12133 · cs.AI · Submitted 2026-08-12 · Read on arXiv

Shivali Dalmia, Sumukha Thoppanahalli, Mohammadreza Sediqin, Abhishek Mukherji

Centific Research

cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: GUIDE: Governed Unified Intelligence for Document-to-Artifact Generation in Enterprise Settings This paper introduces GUIDE, a governed multi-agent framework designed to transform heterogeneous

Terminology

Summary

GUIDE: Governed Unified Intelligence for Document-to-Artifact Generation in Enterprise Settings

This paper introduces GUIDE, a governed multi-agent framework designed to transform heterogeneous enterprise guideline documents into structured, deployment-ready work artifacts. The problem addressed is that enterprise annotation pipelines rely on unstructured guideline documents that must be converted into structured, executable work artifacts before labeling can begin. Quality managers (QMs) and project managers (PMs) currently perform this manually—reading guidelines, inferring rules, resolving ambiguities, and assembling annotator instructions and statements of work. Each document takes 2–3 days, is error-prone, and must be redone whenever documents are updated.

The paper frames this as a data management challenge: Rules extracted from heterogeneous documents must be versioned, deduplicated, validated against a schema, and made traceable from source to artifact. Enterprise guideline documents compound this because "text is encoded as positional tokens rather than semantic units; tables span pages, contain merged cells, or appear as images; and visual elements such as annotation examples and bounding box diagrams carry semantically critical information not captured in text."

GUIDE is architected around a central versioned rule store: "schema-enforced tables keyed by stable rule ids serving as the shared data layer. Agents read from and write to this store through Pydantic-validated contracts, ensuring no downstream agent ever consumes structurally invalid data." This yields three properties essential for enterprise deployment: provenance (every artifact traces to its source rules and originating document), versioning (rule updates across document revisions are reconciled rather than reprocessed), and auditability (every HITL decision is logged against a stable identifier).

The system comprises six components: Parsing Agent, Rule Extraction Agent, Consistency Module, Evaluation Module, HITL Controller, and Artifact Generation Agent.

The Parsing Agent separates deterministic text extraction from VLM processing. Text is extracted from PDF (PyMuPDF), DOCX (internal XML), and PPTX (LibreOffice → PDF) without language models. Visual content is processed using Qwen2.5-VL after MD5-based image deduplication. Extraction quality is assessed via Q = 1 − (garbage + mojibake + repetition + silent skip) where each term is a normalized defect rate.

The Rule Extraction Agent applies a two-stage pipeline: open-domain extraction identifying candidate rules with source spans and confidence scores, followed by normalization into a fixed 26-field schema. Rule type determines persona routing: evaluation-criteria, edge-case, and qa-process rules route to the QM workbench; worker-requirements and delivery-schema rules route to the PM workbench. Rules pass through the Consistency Module, which applies embedding-based similarity filtering followed by NLI classification for deduplication and version alignment. Gap analysis identifies missing or ambiguous aspects as structured GapObjects; resolved gaps generate ClarificationRecords and new RuleUnits.

The Evaluation Engine applies a two-stage validation pipeline. L1 (Structural) applies deterministic Pydantic validation: 28 constraints for RuleUnit, 4 for ExampleObject, and 8 for GapObject. Failed objects are retried, corrected, or rejected by error severity. L2 (Semantic) scores passing objects across K quality dimensions using an LLM-as-judge with K=5 for RuleUnit (clarity, persona fit, completeness, category fit, severity fit), K=3 for ExampleObject (rule alignment, discriminability, input realism), and K=3 for GapObject (question quality, severity calibration, gap type fit). Routing is by minimum dimension score: min k s k ≥ 4 auto-approves; any s k ∈ 2, 3 routes to HITL; any s k = 1 rejects. Objects with rule source = inferred always route to HITL.

HITL is staged and dependency-aware: Phase 1 reviews RuleUnits; Phase 2 reviews GapObjects conditioned on the approved rule set; Phase 3 reviews ExampleObjects conditioned on finalized rules and resolved gaps.

The Artifact Generation Agent transforms approved rules, resolved gaps, and evaluated examples into eight deployment-ready artifacts via Pydantic-constrained templates: annotator guidelines, QA strategy, QA rubric, reviewer instructions, gaps document, QA agent specification, annotator SOW, and job description/requisition. Each artifact is evaluated via ART = w1·RC + w2·SC + w3·PA + w4·CSC where RC (rule coverage), SC (structural conformance), PA (persona appropriateness via Flesch readability and LLM judgment), and CSC (cross-section contradiction, evaluated by Claude Sonnet 4.6) are combined with empirically tuned weights. ART ≥ 4.0 auto-approves; 3.5 ≤ ART < 4.0 routes to human review; ART < 3.5 triggers regeneration.

The corpus consists of 120 enterprise guideline documents from industrial clients (67 text, 23 speech, 16 multimodal, 8 image, 6 video; PDF/DOCX/PPTX formats). Complexity tiers are defined by modality composition: Low (74 KB, 39–46 min), Moderate (1.5 MB, 46–70 min), and High (4.2 MB, 65–125 min). Qwen2.5-VL-32B was selected for all stages based on VLM benchmark results showing the best balance of evidence rate (77.8%), hallucination (20.2%), throughput (232/355 images), and lowest duplication (14.6%).

Content extraction across 120 documents achieves 96% document success with 99.2% page coverage, 97.1% figure recall, 88.3% table recall, and Q=1.00 on successfully processed documents. Failures are caused by VLM timeouts on image-heavy documents and poorly structured tables.

Rule extraction on 115 documents yields 3,896 RuleUnits with 84.8% evidence rate, 82.6% coverage, and 3.2% hallucination. L1 passes 99.1%; the 0.9% flagged are structurally valid but insufficiently precise for direct execution. L2 auto-approves 71.4% with 28.6% routed to HITL and 0% rejected; most HITL cases arise from incomplete semantic coverage where rules capture the primary case but miss edge conditions. The consistency module identifies 26.7% gaps, 3.0% duplications, and 2.9% contradictions.

Artifact generation produces 812 artifacts. Cross-section contradiction (93.7%) and structural conformance (83.1%) reflect strong logical and structural consistency. Rule coverage (56.9%) and persona appropriateness (59.8%) remain the most challenging dimensions. Only 29.8% of artifacts are auto-approved, with 52.0% routed to human review and 18.2% rejected, revealing two primary failure modes: incomplete rule propagation and persona adaptation gaps.

Comparison against a monolithic one-pass Qwen2.5-VL-32B baseline (without the rule store, L1/L2 validation, consistency module, or HITL routing) shows that removing governance degrades every dimension: hallucination rises from 3.2% to 15.7%, duplication from 3.0% to 10.3%, contradictions from 2.9% to 7.8%, and L1 pass rate falls from 99.1% to 93.2%. A layer-wise ablation isolates each layer: L1/L2 routes 1,114 units (28.6%) to review; the consistency module removes 117 duplicates, routes 113 contradictions, and surfaces 1,040 otherwise-undetected gaps.

The paper validates the LLM judge by having 300 rule-level annotations independently labeled by expert annotators blind to judge outputs; the judge achieves precision 0.941, recall 0.974, F1 0.957, and Cohen's κ = 0.813 against these labels.

Limitations noted include: VLM extraction stability decreases on low-quality scans and borderless or merged-cell tables; persona appropriateness and rule coverage remain the most challenging artifact dimensions; the calibration mechanism relies on accumulating zero-edit HITL approvals over deployment cycles, so scoring stability in early cycles remains limited; and the current evaluation covers English enterprise guidelines only.

The paper concludes that GUIDE reduces end-to-end turnaround from 2–3 days to 40–125 minutes per document while maintaining strong extraction fidelity and consistency, demonstrating how governed multi-agent pipelines serve as a principled foundation for data-aware, human-aligned agentic systems where reliability, traceability, and selective human oversight are first-class design goals.

Improvements for AI systems

Improvements to AI Systems:

  1. Governed Multi-Agent Architecture with Centralized Versioned Data Store: Implement a shared, schema-enforced rule store (e.g., Pydantic-validated tables with stable IDs) as the single source of truth. Agents read/write via strict contracts, ensuring no downstream component consumes invalid data. This enables provenance, incremental versioning (updates reconcile rather than reprocess), and full auditability of every decision.

  2. Hybrid Parsing (Deterministic + VLM): Separate text extraction (using PDF/DOCX/PPTX native parsers) from visual content processing (using a VLM like Qwen2.5-VL) with MD5-based image deduplication. Add a quality metric Q = 1 − (garbage + mojibake + repetition + silent skip) to reject or retry low-quality extractions automatically.

  3. Two-Stage Rule Extraction with Persona-Based Routing: First, perform open-domain extraction with confidence scores and source spans; then normalize into a fixed 26-field schema. Route rules by type (e.g., evaluation-criteria → QM workbench; worker-requirements → PM workbench) to specialized downstream agents, improving relevance and reducing noise.

  4. Consistency Module with Embedding + NLI: Apply embedding-based similarity filtering followed by natural language inference (NLI) classification to deduplicate, detect contradictions, and align versions across document revisions. Automatically surface gaps as structured GapObjects for human resolution.

  5. Two-Level Validation (L1 Structural + L2 Semantic): Use deterministic Pydantic constraints (e.g., 28 for RuleUnit) for structural checks, then LLM-as-judge for semantic scoring across K dimensions (e.g., clarity, completeness, persona fit). Route based on minimum dimension score: ≥4 auto-approve, 2–3 → human-in-the-loop (HITL), 1 → reject. Always route inferred rules to HITL.

  6. Staged, Dependency-Aware HITL: Sequence human review in phases—first rules, then gaps (conditioned on approved rules), then examples (conditioned on finalized rules and gaps)—to minimize context switching and rework.

  7. Artifact Generation with Multi-Metric Quality Scoring: Generate deployment-ready artifacts (guidelines, QA strategy, SOW, etc.) via Pydantic-constrained templates. Score each artifact using ART = w1·RC + w2·SC + w3·PA + w4·CSC (rule coverage, structural conformance, persona appropriateness, cross-section contradiction). Auto-approve ≥4.0, human-review 3.5–4.0, regenerate <3.5.

  8. LLM Judge Calibration via Expert Labels: Validate the judge’s outputs against human expert annotations (e.g., precision 0.941, recall 0.974, κ=0.813) and use zero-edit HITL approvals to continuously recalibrate scoring thresholds over deployment cycles.


What the Improved AI System Can Do:

  • Transform unstructured enterprise documents into structured, executable artifacts in 40–125 minutes (vs. 2–3 days manually) with 96% document success, 99.2% page coverage, and Q=1.0 extraction quality.

  • Maintain high fidelity and consistency: 84.8% evidence rate, 3.2% hallucination (vs. 15.7% without governance), 3.0% duplication (vs. 10.3%), 2.9% contradictions (vs. 7.8%), and 99.1% structural pass rate.

  • Automatically route 28.6% of rules to human review for edge cases and incomplete coverage, while auto-approving 71.4%—reducing human workload while preserving oversight.

  • Detect and resolve gaps proactively: Surface 1,040 otherwise-undetected gaps, remove 117 duplicates, and flag 113 contradictions without human pre-screening.

  • Generate 8 types of deployment-ready artifacts (annotator guidelines, QA rubric, reviewer instructions, SOW, job requisition, etc.) with strong structural conformance (83.1%) and near-zero cross-section contradictions (93.7%).

  • Provide full traceability and auditability: Every rule, gap, and artifact links to its source document and version, enabling compliant updates and regulatory review.

  • Scale across modalities (text, speech, image, video, multimodal) and formats (PDF/DOCX/PPTX) with complexity-aware processing times (39–125 minutes based on document size and modality mix).

  • Adapt to new domains by recalibrating the LLM judge via expert-labeled data and accumulating HITL feedback, improving scoring stability over time.

Abstract

Enterprise guideline documents are heterogeneous and multimodal, combining narrative text, complex tables, and embedded images. Existing LLM and VLM systems face hallucinated content, table structure degradation, and lack governed workflows extending beyond extraction to validation and artifact generation. This leaves enterprises to perform this manually, consuming 2-3 days per document. To address this, we introduce GUIDE, a governed multi-agent framework built on a shared versioned rule store with schema-validated inter-agent contracts and end-to-end provenance tracking. Six specialized agents handle parsing, VLM-driven extraction, consistency checking, evaluation, human-in-the-loop (HITL) escalation, and persona-tailored artifact synthesis. Evaluated on 120 real-world enterprise guideline documents, GUIDE achieves 96% document success, extracts 3,896 rules with 71.4% auto-approved, produces 812 deployment-ready artifacts, and reduces turnaround to 40-125 minutes per document.

Sources

Related papers