Excess Separability: Nuisance-Controlled Residual-Stream Probing for Benchmark Contamination Detection
Florian Braun
cs.CL, cs.LG
Submitted: 2026-08-16
Updated: 2026-08-18
Comments: 22 pages, 11 figures, 7 tables. Code and artefacts: https://github.com/mabushi-lab/residual-stream-contamination-probing
Code: https://github.com/mabushi-lab/residual-stream-contamination-probing
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 50/100
The gist: Excess Separability: Nuisance-Controlled Residual-Stream Probing for Benchmark Contamination Detection Summary This paper addresses the problem of benchmark contamination in large language models,
Terminology
Summary
Excess Separability: Nuisance-Controlled Residual-Stream Probing for Benchmark Contamination Detection
Summary
This paper addresses the problem of benchmark contamination in large language models, where a model's performance on a benchmark is inflated because the model has memorized the test items during pretraining. The authors argue that existing detection methods—n-gram overlap search, likelihood-based membership inference, and canary strings—each have a critical missing precondition: the training corpus, a well-chosen test statistic, or foresight at dataset release, respectively.
The paper proposes a new protocol, Residual-Stream Contamination Probing (RSCP), which reads contamination off a linear probe on the model's internal activations. The central claim is that the natural way to do this does not work, and the paper specifies a corrected protocol that survives measurement.
Key methodological contributions and findings:
-
Exchangeability requirement (Requirement E): The reference set R must be constructed so that, under the assumption of no contamination, an analyst told only the text of an item could not do better than chance at saying whether it came from the suspect set S or R. Three constructions satisfy this: E1 (randomised injection), E2 (commissioned twin), and E3 (within-corpus held-out split). Temporal and cross-source splits fail this requirement.
-
Prefix rule: Activations are extracted at the item prefix, never past the answer. This prevents the probe from reading the model's own competence on the item, which would be a confound and a re-description of the accuracy gap.
-
Depth contrast statistic: The tested statistic is a zero-sum contrast on the depth profile of probe accuracy, recentred on a level-matched placebo baseline. This cancels the analyst's control set algebraically, unlike a level-based statistic.
-
Placebo baseline: The baseline is built by splitting the reference set on a surface variable, with the split chosen so its embedding-layer separability matches the observed value. This removes the assumption that surface decodability is flat in depth, which the paper shows fails in both directions.
-
Permutation null on a label-free smoother: The null refits the probe by permuting labels against a cross-fitted smoother, capturing the variance of the probe fit. This is computationally cheap (one matrix-vector product per draw) and correct, unlike an item bootstrap.
-
Validation study: Under two deliberately dissimilar simulators (Sim-A, independent layers; Sim-B, a residual stream with pathologies), the corrected protocol holds the nominal false positive rate where simpler alternatives fail. The level-based statistic's false positive rate tracks the control set's dimension (from 0.03 to 0.99). The uncorrected contrast rejects a true null up to 0.72 of the time when surface decodability rises with depth, and loses all power when it falls. The item bootstrap rejects up to 0.09 of the time where the permutation null holds 0.02. A half-size baseline triples the error rate. The corrected protocol reaches 0.8 power at a separation of roughly one to two accuracy points.
Phase 1 empirical results on real models:
-
Baseline depth profiles are not flat: On a temporal split (WikiMIA), they span up to 29.1 accuracy points and are hump-shaped. On matched splits (Pile train vs. validation), they span 2.7 to 3.9 points.
-
Non-flatness tracks surface difference: Across 6 audits, the correlation between the blind separability (BAnuis) and the baseline span is 0.87.
-
Negative controls behave: All 4 well-matched Pile arms return null, as predicted for a corpus seen approximately once.
-
Protocol refuses a verdict on temporal splits: On WikiMIA, where BAnuis is 0.855–0.856, the protocol refuses a verdict because Requirement E fails.
-
Phase 3 null on deliberately contaminated checkpoints: On Oren et al.'s 1.4B model with PIQA injected at m=50, where exchangeability holds by construction, the contrast is null (Tadj = −0.0040, p = 0.77), even when reading the full record.
Conclusion: The paper's honest summary is that a procedure behaves correctly on synthetic activations where several plausible alternatives do not; it runs on real models; its central assumption was worth removing. Whether a transformer carries a linearly decodable familiarity signal, and at what duplication count, is untouched by any of it, because the only positive result sits behind a failed exchangeability check. Settling that needs the injection experiments of §10 and a benchmark with a reference set that qualifies.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems:
-
Add a contamination-detection layer to model evaluation pipelines. I can implement RSCP as a built-in audit tool that probes internal activations at item prefixes, using the exchangeability requirement (E1–E3) to validate reference sets before any verdict is issued. The improved system can automatically refuse to report benchmark scores when the reference set fails Requirement E (e.g., temporal splits), preventing inflated performance claims.
-
Replace n-gram and likelihood-based contamination checks with depth-contrast probing. I can train a linear probe on residual-stream activations across layers, compute the zero-sum contrast recentered on a placebo baseline, and run the permutation null on a label-free smoother. The improved system can detect contamination at 0.8 power with only 1–2 accuracy points of separation, where current methods miss it or produce false positives up to 0.99.
-
Implement a placebo-baseline generator that matches embedding-layer separability. I can automatically split the reference set on a surface variable (e.g., token frequency, length) to match the observed blind separability, then use that baseline to cancel control-set artifacts. The improved system can avoid the 3× error-rate inflation seen with half-size baselines and the 0.72 false-positive rate when surface decodability rises with depth.
-
Add a label-free permutation null to all probe-based analyses. I can refit the probe by permuting labels against a cross-fitted smoother, using one matrix-vector product per draw. The improved system can hold false positive rates at 0.02 where item bootstraps fail at 0.09, making it safe for high-stakes model releases.
-
Build a refusal mechanism for unverifiable audits. I can encode the rule that if the reference set fails exchangeability (e.g., WikiMIA with BAnuis > 0.85), the system outputs
no verdict
rather than a contamination score. The improved system can transparently flag when its own assumptions are violated, preventing misleading conclusions. -
Enable injection-based validation for future benchmarks. I can design the system to accept deliberately contaminated checkpoints (e.g., PIQA injected at m=50) as calibration inputs, verifying that the contrast returns null (p > 0.05) when exchangeability holds by construction. The improved system can self-test its detection pipeline before deployment on real data.
-
Add a depth-profile diagnostic for surface decodability. I can compute the baseline span (e.g., 2.7–3.9 points on matched splits, up to 29.1 on temporal splits) and correlate it with blind separability (r=0.87). The improved system can warn users when non-flatness is driven by surface differences, prompting them to re-collect the reference set rather than trusting a level-based statistic.
Abstract
Benchmark contamination is diagnosed today with n-gram overlap, with likelihood-based membership inference, or with canary strings, and each needs something usually unavailable: the training corpus, a well-chosen test statistic, or foresight at dataset release. A recent alternative reads contamination off a linear probe on internal activations. We show that the natural way to do this does not work, and specify one that survives measurement. The protocol reports a zero-sum contrast on the depth profile of probe accuracy, recentred on a level-matched placebo baseline, tested against a label-permutation null, with the reference set twice the size of the suspect set. Each choice replaces a simpler alternative we measured and rejected. Reporting the level of excess separability rather than its shape makes the false positive rate track the size of the analyst's own control set, from 0.03 to 0.99 under a true null. Contrasting against a flat depth profile fails in both directions, rejecting a true null 0.72 of the time when surface decodability rises with depth and losing all power when it falls. An item bootstrap holds the fitted probe fixed and rejects up to 0.09 of the time where a permutation null that refits it holds 0.02. A half-size baseline triples the error rate. On real transformers, baseline depth profiles are measurably not flat, spanning up to 29.1 accuracy points on a temporal split, and their non-flatness tracks the surface difference between the item sets (correlation 0.87 over 6 audits), so the correction is largest exactly where it is needed. All 4 well-matched Pile arms return null, and the protocol refuses a verdict on the temporal split rather than reporting one. What this does not establish is whether transformers carry a familiarity direction at all: the only positive sits on the split where exchangeability fails. Implementation, tests and audits are released.
Sources
- Understanding intermediate layers using linear classifier probes
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- Blind Baselines Beat Membership Inference Attacks for Foundation Models
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Language Models (Mostly) Know What They Know
- Scalable Extraction of Training Data from (Production) Language Models
- Rethinking Benchmark and Contamination for Language Models with Rephrased Samples
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering