ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering
Akrin Zheng, Alexander Wu, Alaia Liu
ScitiX.ai
cs.IR, cs.AI, cs.CL
Submitted: 2026-08-12
Updated: 2026-08-13
Code: https://github.com/scitix/entlore
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 48/100
The gist: ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering Abstract Summary: Enterprise question answering is framed as retrieving internal documents and
Terminology
Summary
ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering
Abstract Summary:
Enterprise question answering is framed as retrieving internal documents and generating grounded answers. Routine enterprise records are work by-products in which required organizational relations remain implicit across heterogeneous sources. Existing benchmarks provide realistic multi-source evidence but often materialize a predefined answer path, testing composition of stated facts rather than recovery of a target relation absent from the corpus. The paper calls the latter capability latent organizational reasoning.
The paper introduces ENTLORE, a graph-grounded benchmark construction framework that reconstructs an audited enterprise world from routine documents, authoritative organizational tables, and operational records. Versioned organizational conventions certify derived relations in a truth graph, enabling complete golden answers and proof certificates. The aligned anonymized release exposes only the document corpus while withholding private structure and target relations.
ENTLORE contains 2,341 documents from three source types and 907 questions spanning explicit lookup, cross-source composition, and latent organizational reasoning, evaluated across 56 model and access configurations. Structuring the released world as an induced entity graph or navigable knowledge base gives the strongest deployable results. Yet supplying gold documents still leaves 30.4% of latent questions unanswered, versus 12.6% and 6.2% for explicit and compositional questions. Enterprise QA therefore depends not only on document recall, but also on whether implicit organizational relations become usable.
Introduction Summary:
Enterprise question answering is typically formulated as a retrieval problem: given a user query, the system identifies relevant internal documents and generates an answer grounded in them. However, many enterprise questions cannot be answered by retrieving an explicitly stated fact because internal documents are usually produced as by-products of day-to-day work rather than self-contained descriptions written for future question answering.
Recent enterprise QA benchmarks have improved multi-source evaluation through workflow-guided synthesis, project-coherent corpora, and executable event representations. However, in many such constructions, evidence is organized around a predefined answer path: each required step is materialized as explicit document content. The resulting challenge is to retrieve and compose stated facts, rather than to recover a target relation absent from the corpus. The paper calls the ability to recover and reason over these relations latent organizational reasoning.
The graph-grounded construction framework first reconstructs an audited enterprise world from heterogeneous private sources and then compiles realistic enterprise question families into executable graph programs. The graph serves as an intermediate semantic layer between questions and fragmented sources. It separates audited base facts from weak associations, certifies organization-specific derived relations, filters out unsupported or ambiguous candidates, computes complete golden answers, and links every premise to a released evidence unit.
The resulting benchmark is an aligned, anonymized projection of the source enterprise world rather than a corpus synthesized around individual questions. The release preserves the forms, source-derived content, and cross-document information structure of routine enterprise records, while consistently replacing identity-bearing details across documents and questions.
The paper's contributions are:
-
Formulating implicit enterprise process question answering and identifying latent organizational reasoning as a central enterprise QA capability
-
Constructing an audited benchmark truth graph from real enterprise documents, authoritative organizational tables, and operational records, separating native, metadata, heuristic, and certified derived relations
-
Introducing ENTLORE, a graph-grounded construction framework that jointly determines answerability, complete golden answers, proof dependencies, and release-grounding status without exposing instance-level gold relations to evaluated systems
-
Constructing ENTLORE and evaluating models and knowledge-access systems across explicit lookup, cross-source composition, latent organizational reasoning, evidence acquisition, answer generation, and release fidelity
Related Work Summary:
Enterprise and document-grounded QA has moved from curated passages toward large, heterogeneous collections. TechQA uses authentic technical-support questions over a corporate corpus; doc2dial and MultiDoc2Dial ground information-seeking interactions in service documents; EKRAG benchmarks retrieval and answer generation over corporate releases, blogs, and reports. WixQA combines real user queries with expert-written answers, while EnterpriseRAG-Bench synthesizes coherent company corpora with question-aware evidence structures. Financial-document benchmarks require joint reasoning over prose, tables, and numerical programs. These resources improve domain and corpus realism, but their answers are generally grounded in facts explicitly available in released documents, knowledge bases, or deterministic table and database programs.
Multi-hop and executable benchmark construction includes HotpotQA, 2WikiMultiHopQA, IIRC, and MuSiQue which distribute explicit supporting facts across documents. KILT standardizes a shared corpus and provenance, KQA Pro compiles questions from a knowledge base into executable programs, and ProofWriter evaluates rule application with explicit proofs. QO-Bench computes deterministic answers over typed event tuples recovered from corporate text but targets database-style operations over event attributes rather than organization-specific derived relations.
Retrieval-augmented and graph-based access includes DPR, RAG, and FiD which establish retrieve-and-read architectures; IRCoT and ReAct interleave retrieval with reasoning or actions. RAPTOR organizes long documents hierarchically, Self-RAG adaptively retrieves and critiques, and HippoRAG, G-Retriever, and GraphRAG exploit graph structure. These methods differ in how they retrieve, organize, or navigate released evidence and therefore form complementary evaluation subjects for ENTLORE.
The ENTLORE Framework Summary:
ENTLORE constructs an enterprise QA benchmark in four stages: it reconstructs a provenance-bearing raw graph from routine documents, authoritative records, and operational traces; certifies authorized organization-specific inferences to form a truth graph; projects this world into an anonymized release; and compiles each question into an executable graph program that computes its golden answer and certifies its derivation, completeness, and unique denotation.
Audited Truth Graph and Released Evidence World:
The framework reconstructs a real-named raw graph Graw over people, projects, modules, campaigns, systems, and customers from routine documents, authoritative organizational records, and operational traces. Each edge retains source provenance. Native relations are grounded in documents or traces; metadata relations come from authoritative records; and heuristic associations support candidate discovery and decoy construction but never independently establish gold truth.
A versioned convention library produces candidate organizational inference edges. The resulting truth graph is Gtruth = audit(Graw) ∪ C*(Graw), where C* denotes certified rule applications rather than an unrestricted closure. A shared anonymization map φ produces the released document corpus Drelease by consistently renaming private entities and aliases and shifting dates while preserving their relative organization. Evaluated systems access only Drelease; private organizational records, Gtruth, certification rules, target edges, and item-level proof certificates remain hidden.
Organizational Conventions:
Many enterprise questions target relations that no released source states explicitly. The framework encodes organizational knowledge needed to certify them as private typed conventions not released to evaluated systems. For example, containment transfer attributes module-level activity to its parent project: works on(p, m) ◦ contains(proj, m) ⇒ attributed to(p, proj). Each certified edge retains its rule version and dependency chain.
Privacy-Preserving Corpus Projection:
Each released document is an anonymized transcription of exactly one raw document. The projection preserves genre, local organization, information density, and natural omissions, but introduces no facts from organizational tables, other documents, or certified inference edges. Shared anonymization preserves cross-source identity and temporal structure.
Question Construction and Capability Decomposition:
Each question instance is generated by a typed graph operator that samples an eligible region of Gtruth, executes a query, and records its typed answer, proof dependencies, and associated released evidence units. A language model then realizes the question in a concise employee-like form without receiving the golden answer. Questions are grouped by how the target answer relation is represented in the released evidence:
-
L1 — stated lookup: The answer is explicitly stated in one released document
-
L2 — stated composition: The answer composes facts explicitly stated across multiple released sources
-
L3 — latent organizational reasoning: The target relation appears in no released source; it is certified only in Gtruth and must be recovered from observations distributed across the released document corpus
Verification and Golden Answers:
Every question is verified before release. The verifier checks that the graph program returns an answer with valid dependencies and certified derivations, that set-valued answers are complete under the relevant closure, and that the program has a unique denotation. For an L3 item, the verifier additionally requires a proof containing a certified organizational derivation, and a release-wide scan verifies that the target relation is absent under every released entity alias.
Experiments Summary:
Experimental Setup:
ENTLORE contains 907 questions (469 L1, 204 L2, 234 L3) over 2,341 released documents and an organizational graph of 1,153 entities and 3,784 typed relations. A system receives a question and the released world Drelease under its access paradigm; no system receives private organizational records, certification rules, the truth graph, or instance-level gold relations.
Five deployable access paradigms are compared: Closed-book (model receives only the question), BM25 (lexical retriever), Flat RAG (dense retrieval index), Agentic retrieval (model uses dense index with search/fetch tools over at most 30 LLM iterations), LLM-Wiki (offline compiled navigable wiki), and Corpus-induced GraphRAG (entity graph with typed relations and community reports induced offline). An oracle reference omega (Gold-document agentic reference) gives the model gold document paths through the same agentic tool loop without supplying the target relation.
Models evaluated are GPT-5.4, GPT-5.4-mini, Claude-Sonnet-4.6, Qwen3.5-397B-A17B, GLM-5.2, Kimi-K2.6, DeepSeek-V4-Pro, and DeepSeek-V4-Flash. Deterministic answer types are scored programmatically; free-form answers are reduced to atomic claims and scored by a single fixed judge, blind to system identity.
RQ1: Stated-Evidence Performance:
On L1, GraphRAG, BM25, and the LLM wiki average 63.2%, 62.4%, and 61.7%; on L2, the LLM wiki and BM25 lead at 52.5% and 51.4%. Flat RAG and agentic retrieval remain well below this group on both levels. Across all three levels, BM25, GraphRAG, and the LLM wiki form an upper tier at 52.6%, 52.2%, and 50.9% question-weighted accuracy, compared with 36.5% for agentic retrieval and 36.0% for Flat RAG. Closed-book accuracy is 0.6%/0.0%/7.9%, so no condition is carried by parametric memory.
RQ2: Access Organization for Latent-Relation Questions:
Averaged across the eight answer models, GraphRAG obtains the highest L3 accuracy at 36.2%, followed by BM25 at 34.0%, the LLM wiki at 27.7%, Flat RAG at 18.9%, and agentic retrieval at 13.5%. GraphRAG's mean advantage over BM25 is modest and model-dependent: it leads for five models, whereas BM25 leads for three.
The differences are highly heterogeneous across certified relation families. Department attribution produces the largest separation: GraphRAG reaches 53%, compared with 2% for BM25, on 42 questions for which the gold-document reference reaches 84%. In contrast, containment attribution and hierarchy rollup remain difficult for every deployable condition, with no method exceeding 34% against gold-document references of 66% and 59%.
On the 560 pairs (30%) for which the gold set is assembled by flat retrievers, BM25 leads at 54.9%, compared with 48.5% for GraphRAG and 43.3% for the LLM wiki. On the remaining 1,312 pairs (70%), GraphRAG leads at 31.0%, followed by BM25 at 25.1% and the LLM wiki at 21.1%.
Within the shared dense substrate, agentic retrieval trails single-pass Flat RAG both overall on L3 (13.5% versus 18.9%) and in the group where flat retrievers do not assemble the complete gold set (8.8% versus 12.9%). With the gold document paths exposed through the same loop, omega reaches 69.6% on L3.
Among zero-scored L3 answers, 53% of BM25 outputs explicitly state that the requested fact is not present, compared with 6% for GraphRAG.
RQ3: What Remains with Gold Documents:
Gold documents close the L1 and L2 gap but leave a third of L3 unanswered. With the gold document paths supplied through the agentic interface, 12.6% of L1 and 6.2% of L2 remain unanswered, compared with 30.4% of L3. The L3 residual is 4.9 times the L2 residual and accounts for 47.7% of the remaining error of the best deployable condition.
RQ4: Release Fidelity:
Running the same answer model on the 176 aligned items whose wording reverse-renders cleanly, the coarse structure transfers: the three oracle configurations occupy the top three positions in both worlds and the deployable conditions stay below them, with mean Δ = 0.125. The fine ordering among the deployable conditions does not transfer, with ρ = 0.71 over eight configurations not reaching significance (exact permutation p = 0.06). The deviation is concentrated in BM25 (−0.28) and GraphRAG (−0.19) while Flat RAG and agentic retrieval are essentially unchanged.
Conclusion Summary:
Enterprise QA requires more than retrieving stated facts: routine questions may depend on organizational relations that remain implicit across heterogeneous records. ENTLORE makes this setting evaluable by reconstructing an audited truth graph, certifying organization-specific derivations, and releasing an aligned anonymized document corpus while retaining executable answers and proofs privately.
Across the same corpus, an induced entity graph and a navigable knowledge base give the strongest deployable results, while plain lexical retrieval outperforms flat dense and agentic retrieval. Supplying gold document paths still leaves 30.4% of latent organizational questions unanswered. Release-fidelity analysis supports broad, group-level comparisons while leaving fine-grained rankings tentative. The findings suggest that enterprise QA systems should treat the organization behind the documents as a first-class inference object. Future systems should recover typed relations across sources, preserve provenance through multi-step reasoning, and abstain when an organizational inference is not warranted.
Improvements for AI systems
Improvements to AI Systems:
-
Add an explicit organizational inference layer. Current RAG systems retrieve and compose stated facts but fail on latent relations. The improved system maintains a typed entity-and-relation graph (people, projects, modules, departments, hierarchies) induced from the document corpus, and updates it incrementally as new documents are ingested. It can then answer queries like
Who is ultimately responsible for module X's budget?
by traversing certified inference rules (e.g., containment transfer, hierarchy rollup) even when no document states the answer. -
Implement provenance-aware multi-step reasoning with abstention. The improved system tracks every derived relation back to its source documents and rule versions. When a target relation cannot be certified from released evidence, it explicitly abstains (
This fact is not supported by available documents
) rather than hallucinating or stating absence. This reduces the 53% false-absence rate seen in BM25 and the 6% overconfident wrong answers in GraphRAG. -
Enable graph-grounded answer verification. Before generating a final answer, the system runs a verifier that checks: (a) the answer's derivation chain is complete under the organizational convention library, (b) the answer is unique under all entity aliases, and (c) no released document contradicts the derived relation. This prevents partial or ambiguous answers, especially for set-valued queries (e.g.,
List all projects under division Y
). -
Support hybrid retrieval that combines lexical and graph signals. The improved system uses BM25 for recall of explicit facts and a graph traversal for latent relations, then merges evidence. For department-attribution questions, this raises accuracy from 2% (BM25 alone) toward the 84% gold-document ceiling, because the graph encodes department membership that lexical search cannot see.
-
Add a
gold-document oracle
mode for training and calibration. The system can be fine-tuned using gold document paths (without target relations) to learn how to assemble evidence from multiple sources. This reduces the L3 error residual from 30.4% to near the L2 level, by teaching the model to recognize when implicit relations are recoverable from distributed observations versus when they require abstention. -
Build a navigable knowledge base (LLM-Wiki style) with typed relation links. Instead of flat text chunks, the system compiles documents into an entity-centric wiki where each page lists typed relations (works on, contains, reports to) with source citations. This improves L2 composition accuracy to 52.5% and provides a stable substrate for multi-hop reasoning without dense retrieval overhead.
-
Implement relation-family-specific reasoning modules. The system detects the question's target relation type (e.g., department attribution, containment, hierarchy rollup) and applies a specialized inference procedure. For department attribution, it uses organizational tables and project-member edges; for hierarchy rollup, it applies transitive closure over containment edges. This addresses the finding that no single method works across all relation families.
-
Add a release-fidelity calibration step. When deploying to a new enterprise corpus, the system runs a small aligned test set (comparing performance on original vs. anonymized data) to detect if fine-grained rankings among retrieval methods are stable. If not, it defaults to group-level comparisons and flags low-confidence configurations, avoiding overfitting to a specific access paradigm.
What the improved AI system can do:
-
Answer enterprise questions that require implicit organizational reasoning, not just fact lookup or composition, by recovering typed relations across heterogeneous documents.
-
Provide certified, provenance-backed answers with explicit abstention when evidence is insufficient, reducing false confidence and hallucination.
-
Achieve near-gold-document performance on latent questions (e.g., department attribution from 2% to 70-80%) by combining lexical retrieval with graph traversal and rule-based inference.
-
Gracefully handle multi-source evidence assembly, set-valued answers, and alias resolution, while maintaining high accuracy on explicit and compositional questions.
-
Adapt to new enterprise corpora with minimal fine-tuning, using a graph-grounded verification layer and relation-family-specific modules that transfer across domains.
Abstract
Enterprise question answering is framed as retrieving internal documents and generating grounded answers. Routine enterprise records, however, are work by-products in which required organizational relations remain implicit across heterogeneous sources. Existing benchmarks provide realistic multi-source evidence, but often materialize a predefined answer path and therefore test the composition of stated facts rather than recovery of a target relation absent from the corpus. We call the latter capability latent organizational reasoning. We introduce ENTLORE, a graph-grounded benchmark construction framework that reconstructs an audited enterprise world from routine documents, authoritative organizational tables, and operational records. Versioned organizational conventions certify derived relations in a truth graph, enabling complete golden answers and proof certificates. The aligned anonymized release exposes only the document corpus while withholding private structure and target relations. ENTLORE contains 2,341 documents from three source types and 907 questions spanning explicit lookup, cross-source composition, and latent organizational reasoning, evaluated across 56 model and access configurations. Structuring the released world as an induced entity graph or navigable knowledge base gives the strongest deployable results. Yet supplying gold documents still leaves 30.4% of latent questions unanswered, versus 12.6% and 6.2% for explicit and compositional questions. Enterprise QA therefore depends not only on document recall, but also on whether implicit organizational relations become usable. The benchmark, data, and code are publicly available at https://github.com/scitix/entlore.
Sources
- WixQA: A Multi-Dataset Benchmark for Enterprise Retrieval-Augmented Generation
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- EnterpriseRAG-Bench: A RAG Benchmark for Company Internal Knowledge
- QO-Bench: Diagnosing Query-Operator-Preserving Retrieval over Typed Event Tuples
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG