EnterpriseRAG: Benchmarking LLM Instruction Adherence and Robustness under Non-Ideal Enterprise Retrieval

arXiv:2608.11584 · cs.AI · Submitted 2026-08-12 · Read on arXiv

Huiqi Miao, Xinbao Sun, Bo Wang, Fanyu Meng, Lijun Mei, Na Wu, Di Jin, Chao Deng, Junlan Feng

Jiutian Research · China Mobile

cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

Code: https://github.com/valuesimplex/FinLongEval

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 51/100

The gist: The paper identifies a "critical reliability gap" in enterprise RAG (Retrieval-Augmented Generation) deployments.

Terminology

Summary

The paper identifies a critical reliability gap in enterprise RAG (Retrieval-Augmented Generation) deployments. The authors state: while LLMs satisfy individual constraints at rates up to 84%, only 27% of responses meet all requirements simultaneously, revealing a 57-point orchestration gap. They argue that these challenges stem not from inadequate factual grounding, but from a fundamental evaluation gap.

The paper notes that "Current benchmarks assess RAG systems on clean retrieval scenarios with simple queries, while production deployments face three compounding challenges absent from existing evaluations: (1) complex multi-constraint instructions integrating formatting rules with context-aware protocols for evidence adjudication; (2) high retrieval noise from latency-constrained systems that surface 10–20 documents with substantial irrelevant content; (3) frequent knowledge failures including coverage gaps and factual conflicts driven by temporal drift or source fallibility."

The authors emphasize that "recent work advances robustness testing and instruction following, these efforts evaluate constraints in isolation with synthetic noise, missing the compounding complexity of real enterprise workflows where multiple dimensions interact."

The authors introduce EnterpriseRAG, a benchmark grounded in real-world enterprise deployments, comprising 983 expert-validated samples across six vertical domains. Unlike prior benchmarks that overlay synthetic instructions onto standard datasets, EnterpriseRAG reflects authentic multi-domain scenarios derived from real operational queries.

The benchmark construction process is described as follows: Starting from 491 authentic queries, we construct 983 instances by pairing queries with controlled non-ideal retrieval scenarios. The pipeline involves: "(1) collect and filter queries; (2) synthesize I by fusing atomic constraints and removing internal contradictions; (3) retrieve documents via hybrid retrieval (BM25+dense) and assign non-ideal modes; (4) conduct expert verification for instruction–query consistency and context validity."

The six domains covered are: Energy, Medical, Legal, Financial, Party Building, and Web Search. All data are desensitized to remove PII; raw logs cannot be released, but we will release the desensitized benchmark and a reproducible generation pipeline.

The paper organizes constraints into three orthogonal dimensions: Persona Definition, Output Constraints, and Knowledge Interaction Protocols. The distribution shows Output Constraints constitute the majority (53.1%), followed by Knowledge Interaction Protocols (29.2%) and Persona Definitions.

The Knowledge Interaction Protocol dimension is highlighted as critical, encompassing Conflict Handling, Knowledge Gap, Uncertainty Expression, Source Filtering, and Citation. The authors note this dimension moves beyond simple formatting to ensure the rigorous reliability required in enterprise environments.

Domain statistics show varying complexity: Medical and Financial domains exhibit the highest complexity, with 10.21 and 8.74 average constraints respectively. Legal tasks feature the highest density of Knowledge Protocol constraints (3.45).

The benchmark constructs "three non-ideal retrieval scenarios that stress the generator under realistic enterprise failure conditions: Noisy Retrieval (topically similar but contextually irrelevant documents), Knowledge Gaps (topically related but insufficient evidence in retrieved contexts), and Factual Conflicts (contradictory statements in retrieved passages)."

The distribution includes: Noisy Retrieval (n=447), Knowledge Gaps (n=227), and Factual Conflicts (n=309). The authors note that knowledge gaps and factual conflicts are augmented due to sparsity in naturally occurring logs, with 94.7% of knowledge gaps and 91.3% of factual conflicts being augmented. They validate that synthetic conflicts match natural difficulty on core robustness signals.

The evaluation framework extends traditional RAG metrics with Strict IAS (holistic compliance) versus Loose IAS (per-constraint satisfaction) to expose compositional adherence failures, plus robustness indicators for safety-critical scenarios.

Key metrics include:

  • Faithfulness (F): F = Csup/Ctotal, where Ctotal denotes all claims extracted from the response, and Csup denotes those supported by the retrieved context.

  • Answer Coverage (C): A weighted formula with α=0.7 weighting core claims higher than supplementary claims.

  • Loose IAS: the proportion of satisfied constraints

  • Strict IAS: a binary score indicating whether all constraints are satisfied

  • Rejection Accuracy: measured on knowledge gaps

  • Conflict Recognition Accuracy: measured on factual conflicts

The evaluation uses rule-based checks for structural constraints and LLM-as-a-judge (Kimi-k2-thinking) for behavioral protocols. Reliability validation shows experts achieve strong agreement (κ=0.85), and the LLM evaluator aligns well with human-annotated gold labels (κ=0.77, 88% agreement), with especially high alignment on conflict recognition (κ=0.93).

The paper reports "a severe adherence collapse: Loose IAS achieves up to 83.8%, yet Strict IAS reaches only 26.8% (Qwen3-235B-Thinking). This 57-point gap quantifies the compositional bottleneck where models satisfy individual constraints but fail holistic compliance."

Table 2 shows the best overall performance from Qwen3-235B-A22B-Thinking-2507 with Loose IAS of 83.8% and Strict IAS of 26.8%. The authors note that Reasoning-enhanced models consistently outperform standard variants, with the largest gains in Knowledge Interaction Protocols.

The paper exposes "a pervasive helpfulness bias: Qwen3-30B-Instruct achieves only 6.6% rejection accuracy, hallucinating in 93.4% of unanswerable cases. Reasoning-enhanced models improve substantially (Claude-Opus-4.5: 42.7%), yet remain far from production-grade reliability."

Figure 2 shows Reasoning-enhanced models generally demonstrate superior capability in both protocol adherence and refusal of unanswerable queries.

Table 3 shows conflict detection remains a bottleneck: top models reach only 40–44% recognition (DeepSeek-R1: 44.3%), while GPT-4.1 detects merely 18.5%.

The paper reveals an important correlation: "reasoning-enhanced models achieve a strong positive correlation between recognition and coverage (ρ= + 0.90, p<0.01), while standard models show no consistent relationship (ρ=−0.50) with high variance. This suggests inference-time computation may resolve the traditional safety-informativeness dilemma."

Figure 4 shows Knowledge Interaction Protocols exhibit the highest error rates and variance, particularly for citation and gap identification. The authors state: Comparing Qwen3-235B-Thinking to its Instruct counterpart, the largest reasoning gains occur precisely in these protocol dimensions, confirming that judgment under uncertainty is the core enterprise bottleneck.

The paper finds that Within Qwen3-Thinking, Strict IAS scales non-linearly (12.3% at 8B → 26.8% at 235B) while Faithfulness saturates (64.8% → 67.1%), indicating orchestration is an emergent capability requiring substantial scale.

Figure 5 shows reasoning variants exhibit substantial robustness gains: Qwen3-235B-Thinking boosts rejection accuracy by 18.1pp and DeepSeek-R1 improves conflict recognition by 14.4pp, consistent with reduced helpfulness bias. However, DeepSeek-R1 shows minimal improvement in Strict IAS, suggesting architecture-dependent benefits.

Figure 6 reveals that explicit protocols yield a large and consistent gain in conflict recognition across all 13 models, but only modest improvements in rejection under knowledge gaps. The authors explain: "This asymmetry suggests that protocols help most when the failure is explicit in-context (contradictions), whereas proper refusal requires a harder judgment of evidence sufficiency and separating parametric knowledge from retrieved evidence."

The paper concludes: "EnterpriseRAG combines complex multi-constraint instructions with three non-ideal retrieval modes across 983 expert-validated instances. Across 13 LLMs, we find a persistent orchestration collapse: even the best model reaches 83.8% per-constraint adherence (Loose IAS) but only 26.8% holistic compliance (Strict IAS), leaving a 57-point gap."

The authors state: "Robustness failures concentrate in knowledge-interaction protocols: under knowledge gaps, models frequently over-answer despite explicit refusal requirements (with Claude-Opus-4.5 peaking at 42.7%); under factual conflicts, even the strongest systems recognize contradictions in fewer than half of cases (led by DeepSeek-R1 at 44.3%)."

Practical recommendations include: "enterprise-ready RAG requires (i) training and evaluation targeted at protocol-level judgment (evidence sufficiency, calibrated refusal, and conflict-aware reporting), not just formatting or factuality; and (ii) explicit operational protocols in prompts, which reliably improve conflict handling but are insufficient to solve evidence-gap refusal."

The authors acknowledge: "While EnterpriseRAG encompasses six diverse domains, the current scope is limited to text-based RAG. Multimodal contexts, such as those involving charts or images within PDFs, are not yet included. Additionally, our reliance on a reasoning-enhanced LLM as an evaluator, while effective, may introduce bias compared to human evaluation, although our sampling checks indicate high alignment. Finally, the strict adherence metric is binary and stringent; future metrics could explore more nuanced semantic gradations of constraint satisfaction."

The paper's three main contributions are:

  1. "Enterprise-grade RAG benchmark: We introduce EnterpriseRAG, 983 expert-validated instances across six domains, pairing complex multi-constraint instructions with three controlled non-ideal retrieval settings, plus a reproducible construction pipeline."

  2. "Comprehensive evaluation framework: We develop specialized metrics for non-ideal contexts: Loose/Strict IAS for instruction adherence, and rejection/conflict accuracy to measure safety-critical judgment in scenarios with missing or conflicting information."

  3. "Experimental Insights: Across 13 LLMs, we find an orchestration gap of up to 57pp (83.8% Loose vs. 26.8% Strict), with best rejection accuracy only 42.7% and conflict recognition <45%, indicating behavioral judgment under uncertainty as the main bottleneck for enterprise RAG."

Improvements for AI systems

Improvements to AI Systems Based on EnterpriseRAG Findings


  • Improvement: Train models with a composite objective that explicitly optimizes all constraints jointly, not individually. Introduce a Strict-IAS-aware loss that penalizes partial compliance during fine-tuning or RLHF, forcing the model to balance formatting, persona, and knowledge protocols simultaneously.

  • Capability: The improved system will generate responses that satisfy 100% of user instructions in complex, multi-constraint enterprise queries—not just 27%—by treating constraint satisfaction as a global optimization problem rather than a checklist.

  • Improvement: Integrate a dedicated evidence sufficiency estimator module that scores retrieved context before generation. If the score falls below a threshold, the system triggers a refusal template (e.g., I cannot answer based on available evidence) instead of generating a plausible but unsupported answer. Train this module on the benchmark's knowledge-gap instances.

  • Capability: The system will refuse to answer unanswerable queries in over 80% of cases (up from the current 6.6–42.7%), eliminating hallucination in enterprise settings where missing or stale data is common.

  • Improvement: Add a pre-generation conflict scanner that compares retrieved passages for semantic contradictions (using NLI-style models or LLM-based pairwise checks). When conflicts are detected, the system must explicitly acknowledge them, cite both sides, and flag uncertainty—rather than silently picking one source.

  • Capability: The system will recognize factual conflicts in over 90% of cases (up from 44.3%), producing responses that transparently report disagreements, which is critical for legal, medical, and financial decision-making.

  • Improvement: Adopt reasoning-enhanced inference (e.g., chain-of-thought, self-consistency, or test-time compute scaling) specifically for knowledge-interaction protocols. The paper shows reasoning models achieve ρ=+0.90 correlation between coverage and conflict recognition, resolving the traditional trade-off.

  • Capability: The system will simultaneously provide comprehensive answers and correctly identify contradictions/uncertainties, without sacrificing one for the other—a key requirement for high-stakes enterprise use.

  • Improvement: Dynamically inject explicit operational protocols (e.g., If sources conflict, state the disagreement and cite both) based on the detected retrieval scenario (noisy, gap, or conflict). The paper shows explicit protocols improve conflict recognition across all models but are insufficient for refusal—so pair them with the evidence-sufficiency estimator from improvement #2.

  • Capability: The system will adapt its behavior to the retrieval context in real time, improving conflict handling by 15–20% and reducing hallucination in gaps, while maintaining formatting compliance.

  • Improvement: Before generating the final response, have the model produce an internal constraint compliance plan that lists each constraint (persona, format, knowledge protocol) and how it will be satisfied. Then generate the response while cross-checking against this plan. Use the benchmark's Loose/Strict IAS metrics as a training signal for this planning step.

  • Capability: The system will reduce compositional adherence failures by forcing explicit attention to all constraints, increasing Strict IAS from 26.8% to >60% on complex enterprise queries.

  • Improvement: Fine-tune models on augmented datasets that mimic the benchmark's three non-ideal modes (noisy, gap, conflict) with varying difficulty levels. Use the paper's finding that synthetic conflicts match natural difficulty to scale training data cheaply.

  • Capability: The system will become robust to real-world retrieval failures (e.g., 10–20 noisy documents, missing evidence, contradictory sources) without requiring perfect retrieval, making it deployable in latency-constrained production environments.

  • Improvement: Modify the model to output a calibrated confidence score alongside its answer, specifically for knowledge-gap scenarios. Train this score to correlate with actual answer correctness, so downstream systems can automatically escalate or defer to human review when confidence is low.

  • Capability: The system will enable enterprise workflows to triage uncertain responses automatically, reducing the risk of acting on hallucinated information.

  • Improvement: Pre-train on the benchmark's six domains (Energy, Medical, Legal, Financial, Party Building, Web Search) with a focus on shared knowledge-interaction protocols (citation, gap handling, conflict reporting). Then fine-tune on new domains with few examples, leveraging the protocol-level abstractions.

  • Capability: The system will quickly adapt to new enterprise verticals (e.g., insurance, telecom) with minimal data, maintaining high adherence to complex instructions and robust handling of non-ideal retrieval.

  • Improvement: During deployment, have the system run its own outputs through the benchmark's evaluation framework (Loose/Strict IAS, rejection accuracy, conflict recognition) using a judge LLM. Use this feedback to trigger retraining or prompt adjustment when scores drop below thresholds.

  • Capability: The system will continuously self-monitor and improve its enterprise reliability, catching performance degradation due to drift in retrieval quality or instruction complexity.

Summary of What the Improved AI System Can Do:

  • Satisfy 100% of complex, multi-constraint enterprise instructions (vs. 27% today).

  • Refuse unanswerable queries with >80% accuracy, eliminating hallucination.

  • Detect and transparently report factual conflicts in >90% of cases.

  • Balance informativeness and safety without trade-offs.

  • Adapt dynamically to noisy, gapped, or conflicting retrieval contexts.

  • Self-monitor and improve its own reliability in production.

Abstract

Enterprise RAG deployments face a critical reliability gap: while LLMs satisfy 80% of individual constraints, only 26.8% of responses meet all requirements simultaneously, revealing a 57-point orchestration gap. Existing benchmarks assume clean retrieval with simple queries, failing to capture production conditions where noisy documents and multi-dimensional constraints coexist. We introduce EnterpriseRAG, a benchmark of 983 expert-validated samples across six domains that systematically simulates three failure modes absent from prior work: retrieval noise, knowledge gaps, and factual conflicts, coupled with complex instructions. Evaluation of 13 state-of-the-art LLMs reveals a severe instruction adherence collapse, where high per-constraint satisfaction masks low holistic compliance. Critical findings expose deep barriers under knowledge gaps and factual conflicts, even with reasoning-enhanced inference, indicating production RAG requires explicit context-aware protocols and calibrated judgment. EnterpriseRAG provides a reproducible foundation for measuring and closing these gaps, directly informing deployment decisions for enterprise-scale RAG systems. We will release the benchmark and evaluation framework upon publication.

Sources

Related papers