Co-FactChecker: A Framework for Human-AI Collaborative Claim Verification Using Large Reasoning Models
cs.CL
Submitted: 2026-04-15
Updated: 2026-09-22
Comments: 13 pages, 3 figures, 3 tables. Under review
License: http://creativecommons.org/licenses/by/4.0/
The gist: Professional fact-checkers rely on domain knowledge and deep contextual understanding to verify claims.
Terminology
Abstract
Professional fact-checkers rely on domain knowledge and deep contextual understanding to verify claims. Large language models (LLMs) and large reasoning models (LRMs) lack such grounding and primarily reason from available evidence alone, creating a mismatch between expert-led and fully automated claim verification. To mitigate this gap, we posit human-AI collaboration as a more promising path forward, where expert feedback, grounded in real-world knowledge and domain expertise, guides the model's reasoning. However, existing LRMs are hard to calibrate to natural language feedback, particularly in a multi-turn interaction setup. We propose Co-FactChecker, a framework for human-AI collaborative claim verification. We introduce a new interaction paradigm that treats the model's thinking trace as a shared scratchpad. Co-FactChecker translates expert feedback into trace-edits that introduce targeted modifications to the trace, sidestepping the shortcomings of dialogue-based interaction. We provide theoretical results showing that trace-editing offers advantages over multi-turn dialogue, and our automatic evaluations demonstrate that Co-FactChecker outperforms existing autonomous and human-AI collaboration approaches. Human evaluations further show that Co-FactChecker is preferred over multi-turn dialogue, producing higher quality reasoning and verdicts along with relatively easier to interpret and more useful thinking traces.
Sources
- ReasonIF: Large Reasoning Models Fail to Follow Instructions During Reasoning
- LLMs Get Lost In Multi-Turn Conversation
- ThinkPilot: Steering Reasoning Models via Automated Think-prefixes Optimization
- When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs
- Inference-Time Scaling for Generalist Reward Modeling
- Unraveling Human-AI Teaming: A Review and Outlook
- Can LLMs Automate Fact-Checking Article Writing?
- Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration
- LLM-Based Human-Agent Collaboration and Interaction Systems: A Survey
- Expert Preference-based Evaluation of Automated Related Work Generation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering