SciTrue: Reliable Scientific Claim Validation with Frontier and Open Language Models at the NTCIR SciClaimEval Task

arXiv:2609.00654 · cs.AI, cs.CL · Submitted 2026-09-01 · Read on arXiv

cs.AI, cs.CL

Submitted: 2026-09-01

Updated: 2026-09-01

Comments: To appear in the Proceedings of the 19th NTCIR Conference (NTCIR-19)

Code: https://github.com/taneset/Sci-claim

Project page: https://sciclaimeval.github.io

License: http://creativecommons.org/licenses/by/4.0/

The gist: We describe the SciTrue team's participation in both subtasks of the NTCIR-19 SciClaimEval task, which asks systems to verify scientific claims against the tables and figures of a paper.

Terminology

Abstract

We describe the SciTrue team's participation in both subtasks of the NTCIR-19 SciClaimEval task, which asks systems to verify scientific claims against the tables and figures of a paper. Rather than tuning a single model, we benchmark eleven frontier and open multimodal models under one honest, per-sample protocol and combine them with light, transparent post-processing. On the official, blind test leaderboard (Section), SciTrue placed first by a clear margin in three of the four evidence-category/subtask combinations, and tied for first on the primary metric in the fourth. Three findings explain the result. First, strong instruction-tuned models are already competitive: Claude Opus 4.8 and Gemma-4-31B each exceed the strongest public baseline (o4-mini), and GPT-5.5 and Claude Fable 5 lead both subtasks (97.7 on Subtask 2). Second, the task's pairing structure is the largest lever: a leak-free pair prior that recovers the Supported/Refuted pairing from the claim text alone (a visible field) and assigns Supported to the higher-confidence evidence raises Subtask-1 pair-accuracy from 72.2 to 93.5, far more than any model swap or ensemble weighting. Third, a case-by-case audit finds that most residual errors are visually-undetectable label-mapping swaps or dataset label noise, so measured accuracy understates the true ability and the fixable-by-modeling headroom is small. Controlled fine-tuning, distillation, and agentic consistency-checking support the same conclusions, and we document throughout a measurement leak---label information reaching a system through the packaging of the data rather than its content---in which the released file ordering encodes the label, including one instance that briefly misled our own pipeline.

Sources

Related papers