SAGE: Governed Artifact Generation from Enterprise Guidelines
cs.AI
Submitted: 2026-09-15
Updated: 2026-09-15
License: http://creativecommons.org/licenses/by/4.0/
The gist: Enterprise guideline documents mix narrative text, complex tables, and embedded images, and converting them into structured work artifacts still takes two to three days of manual effort each.
Terminology
Abstract
Enterprise guideline documents mix narrative text, complex tables, and embedded images, and converting them into structured work artifacts still takes two to three days of manual effort each. Current language and vision-language models extract from such documents but offer no governed workflow beyond extraction: no validation, no consistency checking, no traceable artifact generation. We introduce SAGE, a governed multi-stage LLM pipeline organized around a shared versioned rule store with stable identifiers, schema-validated inter-stage contracts, and end-to-end provenance tracking. Extracted rules undergo deterministic structural validation and LLM-based semantic scoring, then a consistency module that removes duplicates, flags contradictions, and surfaces specification gaps; only uncertain or flagged items reach reviewers, while high-confidence outputs are auto-approved. On 120 documents, SAGE cuts turnaround from days to 20-100 minutes, achieving a 96% document-level success rate with 3.2% hallucination, extracting 3,896 rules and producing 812 artifacts ready for human review; without governance, hallucination rises to 15.7%.
Sources
- Position: Early-Stage Quality Assurance in Annotation Pipelines Is More Cost-Effective Than Late-Stage Validation
- The Design of an LLM-powered Unstructured Analytics System
- MonkeyOCR: Document Parsing with a Structure-Recognition-Relation Triplet Paradigm
- Qwen3-VL Technical Report
- Docling: An Efficient Open-Source Toolkit for AI-driven Document Conversion
- Document AI: Benchmarks, Models and Applications
- Unified Structure Generation for Universal Information Extraction
- olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models
- MMDocBench: Benchmarking Large Vision-Language Models for Fine-Grained Visual Document Understanding
- GoLLIE: Annotation Guidelines improve Zero-Shot Information-Extraction
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
- A Survey of Human-in-the-loop for Machine Learning
- ReAct: Synergizing Reasoning and Acting in Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection