Decomposition-Induced Context-Memory Conflict: When Fact-Checking Pipelines Contradict Their Own Source Text

arXiv:2608.10627 · cs.CL · Submitted 2026-08-11 · Read on arXiv

Yu-Feng Yen

Independent Researcher

cs.CL

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: 15 pages, 1 figure

Code: https://github.com/hollis-png/di-cc-decomposition-conflict

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: Decomposition-Induced Context-Memory Conflict: When Fact-Checking Pipelines Contradict Their Own Source Text This paper introduces and characterizes a failure mode in decompose-then-verify pipelines

Terminology

Summary

Decomposition-Induced Context-Memory Conflict: When Fact-Checking Pipelines Contradict Their Own Source Text

This paper introduces and characterizes a failure mode in decompose-then-verify pipelines (including FActScore-style fact-checkers, hallucination detectors, and long-form factuality evaluators), which first split a passage into atomic claims before checking each one. The authors show that decomposition is not a neutral preprocessing step: "When a language model decomposes a passage, it can be induced to substitute its own parametric belief for what the passage actually says, producing a claim that contradicts the source text it was supposed to summarize faithfully." They call this Decomposition-Induced Context-Memory Conflict (DI-CC) and argue it is mechanistically the same phenomenon as classical context-memory conflict, occurring inside a different pipeline stage than prior work has examined.

Problem definition. A claim exhibits parametric injection if it contains content not entailed by the source passage R (judged by an NLI model) and that content is recoverable from the decomposer's own closed-book knowledge (verified via a second NLI pass against a knowledge dump elicited from the model). DI-CC is the contradictory variant where injected content directly contradicts information already present in R; DI-UE (Decomposition-Induced Unsupported Elaboration) is the weaker sibling where injected content fills a position R never addressed, without contradicting it. The mechanistic hypothesis H0 states that a linear probe trained only on classical conflict data (NQ-Swap), never exposed to any decomposition output, should zero-shot separate DI-CC positions from faithful decompositions and from DI-UE positions.

Main mechanistic result. Using Qwen2.5-7B-Instruct as both decomposer and probe host, with a probe trained on 826 activation vectors from NQ-Swap (413 conflict / 413 no-conflict) and frozen, the authors report: "A linear probe trained only on classical context-memory conflict data (NQ-Swap), never exposed to any decomposition output, significantly separates decomposition positions that produce DI-CC from faithful decompositions, and, under most but not all tested configurations, from non-contradictory elaborations as well, at n=64 DI-CC positives (AUC = 0.86–0.88, 95% CI excludes chance, permutation p < 0.0005, the resolution limit of a 2000-permutation test)." Specifically, under variance-ratio layer selection (layer 20): DI-CC vs. legit AUC = 0.881 [0.841, 0.916]; DI-CC vs. DI-UE AUC = 0.863 [0.806, 0.913]. Under accuracy-based selection (layer 22), DI-CC vs. legit AUC = 0.858 [0.815, 0.897] at n=64. The result reproduces the original n=14 pilot's effect size at a larger, same-construction sample (not an independently-constructed replication), with every positive case independently human-verified (13/14 clean at pilot; the one exception was an NLI numerical-precision limitation).

Elicited construction caveat. The primary evidence uses an elicited construction: a decomposition instruction that explicitly authorizes the decomposer to check facts against its own knowledge and correct errors. This instruction is a necessary condition we discovered empirically: without it, DI-CC does not appear at all, even when the source passage contains a tampered detail the decomposer could in principle catch. The authors are explicit that this establishes existence and mechanism but not natural prevalence.

SelfCheckGPT baseline failure. On this same dataset, an existing reference-free baseline, SelfCheckGPT-style self-consistency sampling, fails to detect DI-CC at all (AUC 0.51, chance-level). This failure is mechanistically predictable: DI-CC's defining stability (recoverable, reproducible parametric content) makes it recur consistently across resamples rather than vary, the opposite of the signal self-consistency methods rely on.

Mitigation attempt. Context-aware decoding (CAD), a training-free mitigation from the classical conflict setting, transfers to decomposition and significantly suppresses DI-CC (4.16% to 2.57%, McNemar's p = 0.000041). However, "This comes at a severe cost: 19.4% of decompositions under coreference-heavy conditions fail to parse at all, and 84% of those failures involve the decomposer fabricating a different person's identity rather than merely omitting detail. This is a faithfulness-completeness trade-off the original CAD paper did not report, and we do not consider it deployment-ready as implemented."

Boundaries of the mechanism. The authors characterize scope precisely: "Its natural occurrence rate is too sparse to regress reliably (0.2–0.4% of atomic claims), it does not manifest on naturally-occurring hallucinated text under FActScore's external support criterion, and it requires a minimum model scale to detect: 3B fails to show the signal under both a standard accuracy-based and a generalization-oriented (variance-ratio) layer-selection criterion, a floor pattern that does not depend on which criterion is used, while 7B succeeds under both. The 14B result succeeds under variance-ratio selection but not accuracy-based selection; the criterion was adopted only after seeing accuracy-based selection fail at 14B, so the authors do not treat this magnitude as confirmed evidence the signal increases with scale, and we omit it from this summary."

Cross-family generalization failure. The H0 mechanistic finding did not replicate on Mistral-7B-Instruct-v0.3: DI-CC vs. legit: AUC = 0.397, 95% CI [0.344, 0.449], permutation p = 0.9995; DI-CC vs. DI-UE: AUC = 0.308, 95% CI [0.238, 0.385], p = 1.0000. Both are below chance. Three alternative explanations (statistical power, layer-selection criterion, labeling quality) were checked and ruled out. The authors read this as evidence that the H0 mechanism, as detected by this probing methodology, is at least in part specific to the Qwen2.5 family rather than a general property of instruction-tuned LLM decomposers. A third family (Falcon-7B-Instruct) produced only 10 DI-CC instances, yielding an uninformative AUC estimate.

Practical detection limitation. The probe's usefulness as a fixed-threshold detector is materially weaker than AUCs suggest: Retraining at a realistic, downsampled class ratio (N = 50 seeds) leaves AUC undegraded but collapses fixed-threshold conflict-class recall from 0.808 to 0.247 on average, with high seed-to-seed variance. The authors state they do not claim it translates into a reliable practical detector at realistic prevalence.

Conclusion. The paper's contribution is a three-level characterization: DI-CC is mechanistically real (with the specificity qualification stated explicitly), partially treatable (CAD transfers but with severe side effects), and precisely bounded in scope (low natural rate, absent on FActScore's natural hallucinations, requires minimum scale). The authors stress: "Together, these three levels are the paper's contribution, not a discount applied to a larger claim we would have preferred to make. They form a rigorous accounting of exactly where, how much, and under what conditions this specific failure mode of decompose-then-verify pipelines actually occurs."

Improvements for AI systems

Improvement 1: Add a decomposition-fidelity guardrail to fact-checking pipelines.

The improved AI system inserts a post-decomposition verification step that checks each atomic claim against the source passage using an NLI model before any external fact-checking occurs. If a claim is not entailed by the source, the system flags it as a potential parametric injection and either (a) re-decomposes with a stricter instruction that forbids using parametric knowledge, or (b) discards the claim and marks the passage as unverifiable via decomposition. This prevents DI-CC from silently corrupting downstream factuality scores. The system can now detect and quarantine contradictory claims that would otherwise inflate or distort fact-checking outputs.

Improvement 2: Implement a scale-aware activation probe for early warning of DI-CC.

The improved system uses a frozen linear probe trained on classical context-memory conflict data (e.g., NQ-Swap) and applies it to the decomposer’s internal activations at layer 20 (variance-ratio selection) for models ≥7B parameters. When the probe’s output exceeds a calibrated threshold, the system halts the decomposition step and falls back to a non-decomposing verifier (e.g., direct passage-level NLI or retrieval-augmented generation). This gives the system a zero-shot, training-free early-warning mechanism that detects DI-CC before it propagates, with AUC 0.86–0.88 on Qwen2.5-7B. The system can now preemptively avoid generating contradictory claims in high-stakes fact-checking tasks, even without seeing any decomposition-specific training data.

Improvement 3: Replace self-consistency sampling with a contradiction-aware consistency metric.

The improved system abandons SelfCheckGPT-style self-consistency for DI-CC detection, since DI-CC’s parametric content is stable across resamples (AUC 0.51, chance-level). Instead, it uses a two-pass NLI approach: (a) generate a knowledge dump from the model’s closed-book memory, (b) check each decomposed claim against both the source passage and the knowledge dump. If a claim is entailed by the knowledge dump but contradicts the source, the system classifies it as DI-CC and flags it. This directly targets the mechanism rather than relying on variance, enabling the system to catch contradictions that self-consistency methods systematically miss.

Improvement 4: Apply context-aware decoding (CAD) only with a parse-failure safety net.

The improved system uses CAD to suppress DI-CC (reducing it from 4.16% to 2.57%) but adds a post-hoc parser that detects coreference-heavy failures (19.4% parse failures, 84% identity fabrications). When a parse failure is detected, the system automatically reverts to standard decoding for that passage and logs a warning. This allows the system to benefit from CAD’s suppression effect while avoiding the severe faithfulness-completeness trade-off, making it deployment-ready for passages with low coreference complexity. The system can now reduce DI-CC without risking identity fabrication in downstream outputs.

Improvement 5: Add a model-family and scale gate for DI-CC-sensitive operations.

The improved system checks the underlying model’s family and size before enabling any DI-CC-sensitive features (e.g., decomposition-based fact-checking or the activation probe). For models <7B parameters or non-Qwen families (e.g., Mistral-7B, where the probe fails with AUC 0.397), the system disables decomposition-based verification entirely and falls back to passage-level fact-checking or retrieval-augmented generation. This prevents false confidence in systems where the mechanism is either absent (3B) or not detectable by the probe (Mistral). The system can now reliably choose the correct verification strategy based on known mechanistic boundaries, avoiding both missed detections and spurious alarms.

Improvement 6: Implement a prevalence-aware detector with recall calibration.

The improved system, when used as a practical detector at realistic DI-CC prevalence (0.2–0.4% of atomic claims), retrains the probe with downsampled class ratios and uses a dynamic threshold that adapts to the observed base rate. It also reports a confidence interval around the recall estimate (which drops from 0.808 to 0.247 at realistic prevalence) rather than a single point value. This gives the system an honest, calibrated detector that avoids overclaiming its utility, and it can be used to flag high-risk passages for human review rather than making autonomous decisions. The system can now provide actionable, uncertainty-aware alerts in real-world fact-checking pipelines where DI-CC is rare but consequential.

Abstract

Decompose-then-verify pipelines, including FActScore-style fact-checkers and long-form factuality evaluators, first split a passage into atomic claims before checking each one. Decomposition itself is treated as a neutral preprocessing step. We show it is not: a decomposer can be induced to substitute its own parametric belief for what the source passage says, producing a claim that contradicts the text it was supposed to summarize faithfully. We call this Decomposition-Induced Context-Memory Conflict (DI-CC) and show it is mechanistically the same phenomenon as classical context-memory conflict, occurring inside a different pipeline stage than prior work has examined. A linear probe trained only on classical context-memory conflict data (NQ-Swap), never exposed to any decomposition output, significantly separates decomposition positions that produce DI-CC from faithful decompositions (AUC = 0.86-0.88, permutation p < 0.0005). An existing reference-free baseline, SelfCheckGPT-style self-consistency sampling, fails to detect DI-CC at all (AUC 0.51, chance-level), because DI-CC content is stably recoverable and recurs across resamples, unlike the variability self-consistency methods rely on. Context-aware decoding, a training-free mitigation from the classical setting, transfers to decomposition and suppresses DI-CC, but at a severe cost: many decompositions under coreference-heavy conditions fail to parse, often because the decomposer fabricates a different identity. We do not consider this mitigation deployment-ready. We further characterize the mechanism's boundaries: its natural occurrence rate is too sparss not manifest on naturally-occurring hallucinatedtext, and it requires a minimum model scale to detecablish DI-CC as a real, mechanistically grounded, andpartially treatable failure mode, with a scope we chhan overstate.

Sources

Related papers