Claim-Level Reliability Assessment for Efficient Test-Time Reasoning

arXiv:2608.11994 · cs.AI, cs.CL · Submitted 2026-08-12 · Read on arXiv

Sen Xu, Wei Wang, Shixi Liu, Jixin Min, Yingwei Dai, Zhibin Yin, Yirong Chen, Junlin Zhang

Sina Weibo Inc.

cs.AI, cs.CL

Submitted: 2026-08-12

Updated: 2026-08-13

Code: https://github.com/WeiboAI/CLR

License: http://creativecommons.org/publicdomain/zero/1.0/

Importance score: 95/100

The gist: The paper introduces Claim-Level Reliability Assessment (CLR), a training-free framework for test-time scaling that reallocates compute from additional solution sampling to targeted verification.

Terminology

Summary

The paper introduces Claim-Level Reliability Assessment (CLR), a training-free framework for test-time scaling that reallocates compute from additional solution sampling to targeted verification. The core principle is claim-level falsification, which exploits the asymmetry between solution construction and claim refutation: constructing a valid solution requires a flawless reasoning path, whereas refuting an incorrect claim requires identifying only a single decisive flaw.

CLR operates through a two-stage inference pipeline. In Stage 1, it samples K solution traces, each accompanied by a compact set of M decision-critical claims. These claims encode intermediate conclusions, constraints, decision points, transformations, or evidence linking the problem to the final prediction, filtering out routine tokens that dilute reliability signals. In Stage 2, the same model independently verifies each claim using only the original problem and the extracted claims, actively searching for disconfirming evidence. The verdicts are encoded as binary outcomes (refuted or not refuted), and CLR maps claim-level verdicts into nonlinear trace-level reliability scores using the formula rk = sk M, where sk is the fraction of claims that survive falsification. This nonlinear penalty suppresses error-prone traces, allowing a reliable minority to overturn an incorrect consensus formed by a majority of flawed traces.

The paper evaluates CLR on four LLMs (Gemma-4-12B-it, GPT-OSS-20B, GPT-OSS-120B, Qwen3.5-27B) across four reasoning benchmarks (HMMT25, HMMT26, CMIMC25, Apex-shortlist). Under matched model-call budgets, CLR@32 is compared against Cons@64, with both using 64 model calls. The results show that CLR generally improves accuracy, token efficiency, or both. For example, on HMMT25, CLR@32 raises the accuracy of Gemma-4-12B-it from 76.67% with Cons@64 to 88.75%. On CMIMC25, CLR@32 improves GPT-OSS-20B from 77.50% with Cons@64 to 82.19% while using 37.0% fewer generated tokens, outperforming Pass@1 by 27.15 percentage points. For GPT-OSS-120B, CLR improves accuracy by up to 5.00 percentage points while reducing token consumption by 21.6–23.7%. For Qwen3.5-27B, whose Cons@64 accuracy already exceeds 90% on three benchmarks, CLR matches or improves accuracy by up to 2.60 percentage points, with its largest token reduction of 14.5% occurring on Apex-shortlist.

The paper decomposes the gains of CLR to show that the improvements arise from falsification-based reliability weighting rather than claim prompting itself. Unweighted aggregation over the same 32 candidates shows that the claim prompt alone reduces single-rollout accuracy by 0.65–4.56 percentage points relative to regular sampling, but after claim-level assessment and reweighting, accuracy improves by 4.48–7.01 points over the same unweighted candidates.

The paper also reports rescue rates, which measure the fraction of recoverable consensus failures that CLR overturns. Across the 16 benchmark–budget settings, pooled rescue rates span roughly 16–48% and average about 37%. This demonstrates that CLR can overturn a substantial fraction of erroneous consensus outcomes while preserving reliable minority traces.

An ablation on the number of claims shows that increasing M from 1 to 3 improves accuracy by 3.13–3.79 percentage points across all benchmarks, with further gains on three benchmarks when increasing to M = 5. The main benefit comes from moving beyond a single claim, with task-dependent returns thereafter.

Finally, the paper examines cross-budget accuracy–efficiency scaling, showing that self-consistency often saturates or fluctuates as more solutions are added, while CLR improves more steadily, yielding a more favorable accuracy–compute scaling trajectory. The curves can nevertheless cross at intermediate budgets, and CLR is not uniformly dominant at every operating point. The paper concludes that claim-level falsification can move the accuracy–token frontier outward when additional count-based samples provide noisy or weakly informative support, while offering less headroom when the base consensus is already reliable.

Improvements for AI systems

Improvements to AI Systems:

  1. Implement claim-level falsification as a verification module: Instead of generating multiple full solutions and voting (self-consistency), the AI system will sample K solution traces, extract M decision-critical claims per trace, and independently verify each claim against the original problem. This shifts compute from redundant sampling to targeted refutation, improving accuracy under fixed model-call budgets.

  2. Add nonlinear reliability weighting for candidate selection: The system will compute trace reliability as r k = s k M, where s k is the fraction of claims that survive falsification. This suppresses traces with even a single refuted claim, allowing a small set of highly reliable traces to override an incorrect majority consensus—enabling the system to correct errors that self-consistency would lock in.

  3. Enable compute reallocation for token efficiency: The system will use CLR@K (K samples + verification) instead of Cons@2K, reducing generated tokens by up to 37% while maintaining or improving accuracy. This makes the AI more cost-effective for long reasoning tasks, especially on resource-constrained deployments.

  4. Incorporate claim extraction as a structured intermediate representation: The system will produce compact, decision-critical claims (intermediate conclusions, constraints, transformations, evidence links) rather than raw reasoning traces. This filters out routine tokens, making verification signals denser and more reliable, and enabling the system to audit its own reasoning steps explicitly.

  5. Add a rescue mechanism for consensus failures: The system will detect when a majority of sampled traces are flawed but a reliable minority exists (via claim-level verdicts), then overturn the erroneous consensus. This improves robustness on hard problems where naive majority voting fails, with rescue rates averaging 37% across benchmarks.

  6. Support adaptive claim-count tuning: The system will adjust the number of extracted claims M (e.g., from 1 to 3 to 5) based on task complexity, since increasing M from 1 to 3 yields consistent accuracy gains (3.13–3.79 percentage points), with diminishing returns thereafter. This allows the AI to trade off verification granularity against compute.

  7. Improve scaling behavior for test-time compute: The system will use CLR’s more favorable accuracy–compute scaling trajectory, which improves steadily with additional samples, unlike self-consistency which saturates or fluctuates. This enables predictable performance gains when scaling up inference compute.

  8. Enable cross-model generalization: The system will apply CLR as a training-free, model-agnostic wrapper (validated on Gemma, GPT-OSS, Qwen3.5), so it can be plugged into any LLM without fine-tuning, improving accuracy and efficiency for both small and large models.

What the Improved AI System Can Do:

  • Solve complex mathematical reasoning problems (e.g., HMMT, CMIMC, Apex-shortlist) with higher accuracy than self-consistency at the same compute budget—e.g., raising accuracy from 76.67% to 88.75% on HMMT25 for a 12B model.

  • Overturn incorrect majority answers by identifying and amplifying reliable minority traces, correcting errors that would otherwise persist under standard voting.

  • Reduce token consumption by up to 37% while maintaining or improving accuracy, making it suitable for latency-sensitive or cost-constrained applications.

  • Self-audit its reasoning by explicitly generating and verifying decision-critical claims, providing a transparent, step-by-step reliability score for each candidate answer.

  • Scale test-time compute more effectively, achieving steady accuracy improvements as more samples are added, rather than plateauing or degrading.

  • Operate without retraining, as the framework is training-free and can be applied to any existing LLM, enabling immediate deployment across different model sizes and architectures.

Abstract

We propose claim-level falsification as a principle for test-time scaling and instantiate it through Claim-Level Reliability Assessment (CLR), a training-free framework that reallocates test-time compute from additional solution sampling to targeted verification. Since whole-trace evaluation often obscures decisive errors due to signal dilution from routine tokens, CLR condenses each reasoning trace into a compact set of decision-critical claims, thereby isolating its logical anchors. Furthermore, recognizing the inherent difficulty of generating entirely correct solutions under fixed model capabilities, CLR shifts the focus to semantic falsification. This approach exploits a fundamental asymmetry between solution construction and claim refutation. Constructing a valid solution requires a flawless reasoning path, whereas refuting an incorrect claim requires identifying only a single decisive flaw. This targeted search for negative evidence systematically compresses the survival space of high-confidence incorrect traces, effectively suppressing erroneous consensus via nonlinear reliability scoring. Across four LLMs and four reasoning benchmarks under matched budgets, CLR generally improves upon pass@1 and self-consistency. On GPT-OSS-20B/CMIMC25, for instance, CLR exceeds pass@1 by 27.15 percentage-points and raises self-consistency accuracy from 77.50% to 82.19% with 37.0% fewer tokens.

Sources

Related papers