Beyond Local Accuracy: A Protocol-Level Identifiability Audit for Controlled LLM Reasoning Evaluation

arXiv:2608.13326 · cs.CL · Submitted 2026-08-13 · Read on arXiv

Junhao Luo, Ning Huang, Ziqi Sha, Wenxuan Tang, Wei Deng

Southwestern University of Finance and Economics

cs.CL

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: 15 pages, 9 figures. Ning Huang, Ziqi Sha, and Wenxuan Tang contributed equally as second authors. Wei Deng is the corresponding author

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 96/100

The gist: This paper introduces a protocol-level identifiability audit for LLM evaluation, addressing the problem that "LLM benchmark scores can be precise even when the observation protocol does not identify

Terminology

Summary

This paper introduces a protocol-level identifiability audit for LLM evaluation, addressing the problem that LLM benchmark scores can be precise even when the observation protocol does not identify the behavioral property they are intended to measure. The authors formalize a distinction between two conflated quantities: static correctness (whether the model's answer matches an oracle on fixed inputs) and intervention-response fidelity (whether the model's answer changes in the direction prescribed by a controlled intervention on the input). They argue these are different estimands, and nothing guarantees that a protocol that measures one also measures the other.

The paper casts evaluation design as an identification problem. Let H be a finite class of deterministic behavioral policies, where a policy h maps an input under a representation contract to a response. An observation support O is the set of responses a protocol collects. The theoretical target is the full-support contract-stable selective-response property **τ full: H → 0,1 **, which equals one only when a policy is base-correct, follows the target oracle, remains stable on matched sham, and does so across both readouts and all six mappings.

The central theorem (Theorem 1) states: τ is point-identified on O if and only if for all h, h′ ∈ H, τ(h) ≠ τ(h′) implies h ≢ O h′. The proof shows that if two policies with different τ values are observationally equivalent, any estimator based on O must assign them the same value, so at least one is wrong. Conversely, if all pairs with different τ are inequivalent, a well-defined correct estimator exists. The theorem is deliberately finite and non-statistical: it states a structural condition on the support and the policy class, independent of sample size.

The study uses solver-grounded three-valued reasoning over open-world signed-Horn theories. An independent solver fixes, before any model run, the oracle label (TRUE/FALSE/UNKNOWN) for each intervention world. Models answer under a representation contract, a readout channel (generated choice or candidate scoring), and a label mapping (one of six S3 permutations of TRUE/FALSE/UNKNOWN to A/B/C).

Observations are organized as an Interventional Response Tensor R w,π,r with indices: world w (base, target, sham), label permutation π (six S3 permutations), and readout r (generated choice or candidate scoring). Full support has 3 × 6 × 2 = 36 cells per model–cluster. The target world applies an edit that changes the solver oracle; the sham world applies an edit that preserves it.

The paper implements seven frozen deterministic policies plus one stochastic mixture. The reference ideal semantic updater has τ full = 1; the other six have τ full = 0 (target inertia, any-edit reactor, fixed-label responder, mapping-conditional shortcut, generated-only failure, scoring-only success).

The complete 2 4 = 16 subset lattice over four support components (target, sham, paired readout, full S3 mapping) is enumerated. Key findings:

  • Base-only support (O0) collapses all seven policies into one observation-equivalence class with 6 cross-estimand collisions (every τ=0 alternative is observationally equivalent to the ideal updater on base alone).

  • Each added axis splits classes: O1 yields 2 classes (3 collisions), O2 yields 3 (2 remain), O3 yields 5 (1 remains), and full support O4 yields 7 singleton classes with 0 collisions (point identified=1).

  • Leave-one-out audit: Every reduced support retains at least one cross-estimand collision. Removing target yields 3 collisions; removing sham, readout, or mapping each yields 1 exclusive witness. Thus all four components are necessary for point identification.

Four corollaries establish axis insufficiency:

  • Base insufficiency: Cannot separate a selective updater from target inertia.

  • Sham insufficiency: Equates a selective updater with an any-edit reactor.

  • Readout insufficiency: Equates a semantic failure with a contract-specific reportability failure.

  • Mapping insufficiency: Equates a semantic updater with an unseen-mapping label shortcut.

The collision structure enables protocol synthesis: finding the smallest support that still separates every cross-estimand pair. For each pair with different τ, the distinguishing cells are collected, and a support is identifying iff it is a hitting set of these distinguishing-cell sets. On the frozen seven-policy class, the result is striking: O = 2 cells—not 36*. There are 26 distinct 2-cell supports (no cell is mandatory): every minimum pairs one target-generated-choice cell with one sham-candidate-scoring cell. Under equal cell costs the minimal cost is 2; if candidate scoring costs three times a generated choice, every 2-cell support costs 4. The authors emphasize: The synthesis procedure is the methodological contribution; its numerical outcome is case-specific.

The diagnostic case spans 24 clusters, six mappings, two readouts, three worlds, and two models (Qwen/Qwen2.5-7B-Instruct and meta-llama/Llama-3.1-8B-Instruct), giving 1,728 cells and 576 paired base/target/sham units.

Key empirical results:

  1. Estimand-separated results: Base accuracy is 0.403 (576 units); local response base-correct is 0.138 (95% CI [0.068, 0.218], 32/232 units); target-follow base-correct is 0.168; sham-stay base-correct is 0.918; local paired response is 0.056 (95% CI [0.026, 0.090]); full-support response is 0.000 (0/48 model–clusters). The headline is conditional: only 0.138 of base-correct units satisfy it... Target following is 0.168, versus 0.918 sham stability, locating the shortfall mainly in target response.

  2. Behavioral composition: A mutually exclusive audit finds 183 reportability failures, 179 base-knowledge failures, and 164 target-inertia cases. Target inertia is therefore the largest base-correct semantic non-response, not the dominant category overall. Sham stability holds for 349/576 units.

  3. Contract and transition robustness: Equal-validity contracts give both constrained-generation variants pair-validity 1.0; pooled candidate scoring has pair-validity 0.9826. Balanced transitions across six directions and 48 clusters show equal-weighted selective response of 0.324 (95% CI [0.304, 0.345]) versus base accuracy 0.620 ([0.600, 0.642]), with P(∆ > 0) = 1.0 over synchronized cluster bootstraps. A second deterministic source shows rates 0.331 ([0.312, 0.350]) and 0.646 ([0.618, 0.674]).

  4. Mapping and readout: Semantic equivariance is 0.0625 (18/288 groups), rising to 0.1023 after validity filtering; fixed-label attachment is 0.0590. Generated-choice validity is 0.4560 versus 0.9942 for scoring. When both readouts are valid, semantic answers always agree (base 132/132, target 125/125, sham 137/137).

  5. Oracle-transition asymmetry: Selective response is more than four times more frequent after TRUE base oracles (0.095) than after FALSE base oracles (0.022), suggesting models update more readily when a previously entailed proposition becomes undetermined.

The authors acknowledge: The structural results are relative to a frozen seven-policy class, not an exhaustive model of natural-language behavior; additional policies may require more support. The setting is a controlled testbed for protocol diagnosis rather than evidence of unrestricted transport to natural-language evaluation. Target and sham differ in edit type and surface form, so their contrast supports protocol discrimination, not a pure edit effect. The stochastic mixture is recovered only at its declared granularity—the analyzer identifies the equivalence class containing the mixture... but does not identify the exact mixture parameters from behavioral observation alone.

The paper's central claim is structural and protocol-relative: identification holds on the frozen policy class and support, and its measurement consequences hold across readout contracts, transition directions, and two deterministic sources. The authors conclude: Evaluation should state its behavioral estimand and test whether support identifies it before interpreting a score. The present study establishes this workflow in a controlled solver-grounded reasoning setting. The audit requires zero model calls and demonstrates how evaluation-design validity can be checked structurally before model inference and why base correctness does not determine intervention-response fidelity.

Improvements for AI systems

Based on this paper, I can improve AI systems in the following specific ways:

1. Add protocol-level identifiability audits to evaluation pipelines. Before running any model inference, I can check whether the observation protocol (which inputs, readouts, and label mappings are collected) uniquely identifies the behavioral property being measured. If two policies with different target properties are observationally equivalent under the protocol, I flag the evaluation as non-identifying and recommend adding missing support components (e.g., sham worlds, paired readouts, or full label mappings).

2. Separate static correctness from intervention-response fidelity in reporting. I will report distinct estimands: (a) base accuracy on fixed inputs, (b) target-following rate given base-correctness, (c) sham-stability rate given base-correctness, and (d) full-support selective response. This prevents conflating model knows the answer with model updates correctly under intervention, which the paper shows can differ dramatically (0.403 base accuracy vs. 0.138 local response fidelity).

3. Synthesize minimal identifying supports to reduce evaluation cost. Given a policy class of interest, I can compute the minimum set of observation cells needed to separate all cross-estimand pairs. This allows me to design evaluations with far fewer cells (e.g., 2 cells instead of 36 in the paper's case) while preserving identifiability, reducing compute and annotation costs without sacrificing validity.

4. Detect and diagnose specific failure modes via behavioral composition audits. I can classify each model response into mutually exclusive categories: reportability failures (cannot produce valid output under the contract), base-knowledge failures (wrong on base), and target-inertia (correct on base but fails to update under intervention). This gives actionable diagnostics—e.g., if target-inertia dominates, the system needs better intervention handling, not more knowledge.

5. Implement leave-one-out support audits to identify critical protocol components. For any evaluation, I can test which support axes (target, sham, readout, mapping) are necessary by removing each and checking if cross-estimand collisions emerge. This tells system designers which protocol elements are indispensable for valid measurement, preventing silent invalidity from omitted controls.

6. Add transition-direction and contract-robustness checks. I will evaluate models across multiple intervention directions (TRUE→UNKNOWN, FALSE→UNKNOWN, etc.) and multiple response contracts (generated choice vs. candidate scoring) to ensure measured behavior is not an artifact of a single surface form. The paper shows selective response varies 4x by base oracle type, so I will report direction-specific rates rather than pooled averages.

7. Build an early-warning system for label-mapping shortcuts. I can test whether model performance depends on the specific permutation mapping TRUE/FALSE/UNKNOWN to A/B/C. If semantic equivariance is low (e.g., 0.0625 in the paper), I flag that the model may be exploiting label-position shortcuts rather than true semantic reasoning, and I will require multi-mapping evaluation for any deployment claim.

8. Use the zero-model-call audit as a pre-training gate. Before spending compute on fine-tuning or prompting, I can run the structural identifiability check on the proposed evaluation design. If the design is non-identifying, I will refuse to run the experiment and instead propose a corrected protocol—saving significant resources and preventing invalid conclusions.

Sources

Related papers