SciTrue: Reliable Scientific Claim Validation with Frontier and Open Language Models at the NTCIR SciClaimEval Task
cs.AI, cs.CL
Submitted: 2026-09-01
Updated: 2026-09-01
Comments: To appear in the Proceedings of the 19th NTCIR Conference (NTCIR-19)
Code: https://github.com/taneset/Sci-claim
Project page: https://sciclaimeval.github.io
License: http://creativecommons.org/licenses/by/4.0/
The gist: We describe the SciTrue team's participation in both subtasks of the NTCIR-19 SciClaimEval task, which asks systems to verify scientific claims against the tables and figures of a paper.
Terminology
Abstract
We describe the SciTrue team's participation in both subtasks of the NTCIR-19 SciClaimEval task, which asks systems to verify scientific claims against the tables and figures of a paper. Rather than tuning a single model, we benchmark eleven frontier and open multimodal models under one honest, per-sample protocol and combine them with light, transparent post-processing. On the official, blind test leaderboard (Section), SciTrue placed first by a clear margin in three of the four evidence-category/subtask combinations, and tied for first on the primary metric in the fourth. Three findings explain the result. First, strong instruction-tuned models are already competitive: Claude Opus 4.8 and Gemma-4-31B each exceed the strongest public baseline (o4-mini), and GPT-5.5 and Claude Fable 5 lead both subtasks (97.7 on Subtask 2). Second, the task's pairing structure is the largest lever: a leak-free pair prior that recovers the Supported/Refuted pairing from the claim text alone (a visible field) and assigns Supported to the higher-confidence evidence raises Subtask-1 pair-accuracy from 72.2 to 93.5, far more than any model swap or ensemble weighting. Third, a case-by-case audit finds that most residual errors are visually-undetectable label-mapping swaps or dataset label noise, so measured accuracy understates the true ability and the fixable-by-modeling headroom is small. Controlled fine-tuning, distillation, and agentic consistency-checking support the same conclusions, and we document throughout a measurement leak---label information reaching a system through the packaging of the data rather than its content---in which the released file ordering encodes the label, including one instance that briefly misled our own pipeline.
Sources
- Qwen2.5-VL Technical Report
- Conflict-Aware Fusion: Mitigating Logic Inertia in Large Language Models via Structured Cognitive Priors
- The Llama 3 Herd of Models
- Gemma: Open Models Based on Gemini Research and Technology
- Distilling the Knowledge in a Neural Network
- GPT-4 Technical Report
- SciClaimEval: Cross-modal Claim Verification in Scientific Papers
- Input-length-shortening and text generation via attention values
- MFC-Bench: Benchmarking Multimodal Fact-Checking with Large Vision-Language Models
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- CoRA: Optimizing Low-Rank Adaptation with Common Subspace of Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection