ReflectFact: Self-Reflective Agents for Improving Comprehension and Reasoning in Multi-Hop Fact Verification

arXiv:2608.12877 · cs.AI · Submitted 2026-08-13 · Read on arXiv

Runze Zhao, Zixin Tang, Xiaoshuai Hao, Leyuan Chang, Xiaopeng Fu, Boyu Qiao, Dongyang Zhang

Zhongguancun Laboratory · Institute of Information Engineering, Chinese Academy of Sciences · Xiaomi EV

cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: 9 pages, 4 figures, 3 tables

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: ReflectFact is a novel self-reflective agent framework for multi-hop fact verification, proposed to address two critical limitations in existing agent-based methods: objective conflicts and knowledge

Terminology

Summary

ReflectFact is a novel self-reflective agent framework for multi-hop fact verification, proposed to address two critical limitations in existing agent-based methods: objective conflicts and knowledge conflicts. Objective conflicts arise when agents performing individual subtasks lack sufficient awareness of the global verification objective, causing their reasoning to deviate from the intended direction. Knowledge conflicts occur when parametric knowledge in an agent is inconsistent with evidential knowledge, referred to as evidence drift, which can lead to unsupported modifications to an otherwise valid claim.

ReflectFact introduces three key tasks coordinated across verification stages. Explicit Reasoning Path Planning builds an evidence-grounded reasoning path through three automated pipelines: Implicit Entity Resolution identifies implicitly referenced entities, Semantic Decomposition breaks claims into atomic subclaims, and Integrative Logical Reasoning constructs coherent logical chains for verdicts. Evidence-Drift Verification makes the agent re-answer by quoting the supporting evidence when a grounded answer merely echoes its parametric prior, thereby calibrating evidence deviation to ensure grounded comprehension. Reasoning Reflection Verification re-examines each reasoning step and regenerates it once an inconsistency is detected, correcting reasoning flaws such as location bias and replacement bias through a global task perspective. Subsequently, the agent aggregates validated reasoning chains to yield reliable verdicts.

The framework leverages the observation that LLMs exhibit stronger verification than generation capabilities, facilitating self-evaluation of reasoning coherence with the overarching objective. Extensive experiments on HOVER and EX-FEVER demonstrate that ReflectFact consistently outperforms the strongest of ten competitive baselines by 3.32% and 2.78% in overall Macro-F1, achieving state-of-the-art performance with robust scalability and interpretability.

The main contributions are: identifying two key limitations faced by agent-based methods—lack of a global reasoning perspective and over-reliance on parametric knowledge—which render them prone to both objective conflicts and knowledge conflicts; proposing ReflectFact, a novel self-reflective agent framework introducing self-reflective post-hoc verification at each reasoning step to address limitations in subtask-level reasoning quality; and achieving significant performance gains on HOVER and EX-FEVER, outperforming the strongest of ten competitive baselines and demonstrating superior adaptability in handling complex fact verification.

In the methodology, multi-hop fact verification is formulated as determining whether a claim c is supported or refuted through multi-step reasoning over evidence E. Explicit Reasoning Path Planning organizes verification into three stages. Implicit Entity Resolution identifies implicit entities within the claim using three components: Locate identifies the span corresponding to an implicit entity description and generates a targeted question; Search accesses external evidence to derive an answer explicitly containing the identified entity; Replace leverages LLMs to replace the original description with the entity found in the answer. Semantic Decomposition verifies the refined claim by examining each semantic component individually, generating sub-questions and obtaining corresponding answers from external knowledge. Integrative Logical Reasoning combines outputs from preceding stages into a coherent reasoning chain using manually constructed chain-of-thought templates.

Evidence-Drift Verification detects evidence drift by introducing an evidence-free counterpart of the same query, comparing the evidence-grounded answer with the parametric answer. If the two answers converge, the sub-task is flagged as a candidate instance of evidence drift, and the agent is required to re-derive the answer while quoting the specific span of evidence supporting its conclusion, re-anchoring the answer to retrieved evidence rather than the LLM's internal prior.

Reasoning Reflection Verification treats the produced result as an object to be checked, exploiting the observation that verification capabilities of LLMs typically surpass generative abilities. It constructs a verification prompt asking the LLM to check whether the output correctly and consistently follows from the input, prepended with explicit task framing: This is part of a fact-checking task, and any error or inconsistency found must be reported. If inconsistency is detected, the output is regenerated given the flagged inconsistency.

Experiments were conducted on HOVER and EX-FEVER datasets. HOVER is a multi-hop dataset derived from English Wikipedia articles, using the validation set of 4,000 claims requiring evidence from up to four Wikipedia articles. EX-FEVER involves 2-hop and 3-hop reasoning with claims created by summarizing and modifying information from hyperlinked Wikipedia documents, evaluated on the test set of 4,071 claims after removing NEI labels. Ten baselines were used across three categories: Vanilla LLM (FLAN-T5, Qwen3, GPT-4o-mini), Inference augmented models (ScandiNLI, DeBERTaV3-NLI, ProgramFC), and Agent-based Fact Verification (HiSS, Factcheck-GPT, StepByStepFV, BiDeV). Macro-F1 was adopted as the primary evaluation metric.

Results show that ReflectFact achieves the best overall performance on both datasets. On HOVER, ReflectFact achieves 83.33% on 2-hop, 76.91% on 3-hop, 73.74% on 4-hop, and 78.51% overall. On EX-FEVER, it achieves 86.68% on 2-hop, 80.55% on 3-hop, and 83.76% overall. The strongest agent-based baseline BiDeV trails ReflectFact by 3.32% and 4.10% in overall performance on HOVER and EX-FEVER, respectively.

Ablation studies show that removing Evidence-Drift Verification causes a moderate but consistent performance drop that widens as the number of hops increases, while removing Reasoning Reflection Verification leads to a substantially larger degradation, confirming that recasting the agent's own reasoning output as an object to be verified is the primary source of ReflectFact's robustness.

Model generalizability experiments instantiate ReflectFact on GPT-4o-mini and Qwen3-8B backbones, showing consistent Macro-F1 improvements across both backbones and all settings, with advantages becoming more pronounced as reasoning depth increases.

Interpretability analysis using ROUGE metrics on EX-FEVER golden explanations shows that ReflectFact outperforms three comparison models (MDR, BERT, GPT) on average by 5.36% and 7.96% in Rouge-1 and Rouge-L metrics, despite not being explicitly designed for generating explanations.

Error type analysis on 40 randomly sampled examples categorizes errors into three types: logical mistake (10%), factual hallucination (82.5%), and thoughts omission (7.5%). Further analysis of errors in Semantic Decomposition shows that 72.7% of errors are concentrated in hallucinations generated by the LLM itself, while decomposing sub-information enhances logical reasoning but also ignores longer dependencies in the text, resulting in some performance loss.

Improvements for AI systems

Improvements to AI Systems:

  1. Self-Reflective Post-Hoc Verification Module: Add a verification step after each reasoning action where the model re-checks its own output against the input and global task objective. This reduces objective conflicts by ensuring every intermediate step aligns with the final goal, not just local subtask success. The improved system can catch and correct reasoning deviations (e.g., location bias, replacement bias) before they propagate.

  2. Evidence-Drift Calibration Mechanism: Implement a dual-answer comparison (evidence-grounded vs. evidence-free) for any claim supported by retrieved documents. If both answers match, force the model to re-answer while quoting specific evidence spans. This prevents over-reliance on parametric memory, reducing hallucinated modifications to valid claims. The system becomes more trustworthy in knowledge-intensive tasks where evidence and internal knowledge may conflict.

  3. Explicit Reasoning Path Planning with Implicit Entity Resolution: Add a pipeline that identifies implicit entities (via locate-search-replace) before decomposing claims. This enables the system to handle ambiguous or under-specified queries by grounding them in external evidence first. The improved system can resolve coreference and implicit references, improving accuracy on multi-hop questions where entities are not explicitly named.

  4. Semantic Decomposition with Atomic Subclaim Verification: Break claims into atomic subclaims and verify each independently against evidence, then recombine via logical templates. This improves modular reasoning and error isolation, allowing the system to pinpoint which part of a claim is unsupported. It also enhances interpretability by producing step-by-step justifications.

  5. Verification-Over-Generation Exploitation: Use the model’s stronger verification capability to critique its own generative outputs. Instead of only generating answers, the system explicitly prompts itself to check consistency (This is part of a fact-checking task; report any error). This leverages inherent LLM strengths, improving reliability without additional training.

  6. Global-Objective-Aware Reasoning Chains: Integrate manually constructed chain-of-thought templates that force the model to link sub-results back to the overarching claim. This reduces objective conflicts by maintaining a global perspective across hops, improving performance on tasks requiring multi-step logical integration.

What the Improved AI System Can Do:

  • Perform multi-hop fact verification with higher accuracy (e.g., +3.32% Macro-F1 on HOVER) by self-correcting reasoning flaws in real time.

  • Distinguish between parametric knowledge and evidence, avoiding unsupported modifications to claims (reducing knowledge conflicts).

  • Handle ambiguous claims with implicit entities by grounding them in external sources before reasoning.

  • Provide interpretable, step-by-step verification trails that align with human logical expectations, improving user trust.

  • Scale to deeper reasoning chains (4-hop) without performance degradation, as self-reflection becomes more critical with complexity.

  • Generate explanations that are more faithful to evidence (higher ROUGE scores) even when not explicitly trained for explanation tasks.

Sources

Related papers