Located but Not Releasable: Silent Gate Inversion and Bounded Linear Release
Xining Xun
Tsingjiao Information Science (Beijing) Company Limited
cs.CL, cs.LG
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: 16 pages, 5 figures, 5 tables
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: This paper presents a fully preregistered, end-to-end stress test of the detect–localize–release pipeline for intervening on language model representations, using a 25.7M-parameter transformer
Terminology
Summary
This paper presents a fully preregistered, end-to-end stress test of the detect–localize–release pipeline for intervening on language model representations, using a 25.7M-parameter transformer (C1) trained on causal-evidence discrimination where a known suppression phenomenon exists (latent causal structure present but behaviorally unused). The authors submit the complete pipeline—detect, localize, and release—to a preregistered protocol where every threshold, every claim template, and every branch of the decision tree was hashed and archived before any corresponding data existed.
The paper's central contribution is a failure decomposition into two dissociable parts. Three main findings emerge:
Finding 1: Localization succeeds. The suppression is site-specific. Using transplant interventions at observation-evidence channels of mid layers, the authors restore target behavior on suppressed worlds. The paired release advantages are 0.563 and 0.854 with 97.5% confidence intervals excluding zero, and the best-site release rate is 0.889. Specifically, L23 pair: 0.5628 [0.3932, 0.7302], L345 pair: 0.8542 [0.6088, 1.0995] (n = 48),
and the best site, L23obs, attains mean release rate Rsite = 0.8886 ≥ 0.5 (C2, pass).
The preregistered claim G1 is invoked: the suppression-by-default under OOD is localized to the observation-evidence row channels, and transplant intervention is sufficient to restore the target behavior.
Finding 2: Gating fails out of distribution (Failure 1). The detector calibrated to trigger on zero out-of-distribution calibration worlds triggers on 6.9–7.3% of held-out in-distribution generations and on zero of the 2,400 held-out generations that actually need it. The paper states: Across all 2,400 CONF50 generations—the worlds for which the entire apparatus exists—the gate triggered zero times: GATE ≡ BASE, and the intervention never happened.
This is described as a complete inversion that silently reduces the gated pipeline to its base model.
The detector learned 'looks like calibration OOD' rather than 'needs intervention.'
Three of five frozen criteria failed: release intercept b = 0.3820 [0.3396, 0.4230] against threshold ≤ 0.08; ID trigger rates 0.0688–0.0734 against cap 0.02; and McNemar test degenerate (b=c=0, p=NaN) with 25/27 negative-ATE worlds sign-reversed against a cap of 14/27.
Finding 3: Linear release is capped (Failure 2). Removing the gate entirely and injecting a per-instance linear direction unconditionally yields a monotone dose–response that plateaus far below the preregistered release margin. The intercept falls monotonically with dose: 0.3820 → 0.3106 → 0.2637
against the frozen release threshold b ≤ 0.08. The full-strength arm remains more than 0.18 above the frozen release threshold.
Per-instance adaptivity adds less than ±0.03 over the fixed direction: +0.0268
and −0.0199
per world. The preregistered readout is R2 NO MARGIN, with the verbatim claim: Per-instance linear directions at this site and resolution likewise cannot be released—the linear release family, fixed and adaptive alike, is bounded as a whole.
The paper concludes: "The failure is therefore doubly located: the detector is OOD-inverted, and the entire family of linear release directions at this site and resolution is bounded away from sufficiency. The two failures are dissociable, and neither overturns the localization result—representation-level localization and behavior-level release come apart."
The paper also reports three stage-1 gates that all failed before any release attempt: behavior-coupled readout does not exist (only 11/50 worlds measurable), representations do not align across evidence types (CKA 0.061–0.068 vs. threshold 0.7), and the answer is detectable but not usable by readout (AUC 0.962 for family separation but worst leave-one-family-out AUC 0.366).
The discussion emphasizes that the representation–behavior gap is not one gap
but decomposes into at least three independent obstacles, each sufficient to sink a naive pipeline. The paper recommends gate-level trigger counts on the target distribution be reported as a first-class metric wherever a gate exists, and notes that "if linear directions are insufficient even where localization is certain, then steering failures elsewhere cannot be automatically attributed to 'wrong layer' or 'wrong direction': the linear parametrization itself is a candidate bottleneck."
Limitations acknowledged include: a single 25.7M model on a synthetic causal domain, the release family tested is linear/single-site/residual-additive, the denominator guard excludes 39/50 stage-1 worlds from rate statistics, the eligibility rate for localization is 7.9% (48/608 worlds), and the gate was calibrated without held-out examples of the positive class's test-time realization. Total measured cost of terminal experiments: 4.22 GPU-hours on a single consumer-class GPU.
Improvements for AI systems
Improvement 1: Add distribution-aware gating with target-distribution trigger validation.
The improved AI system will include a gate that is calibrated and validated on the actual target distribution (e.g., held-out generations that require intervention), not just on calibration OOD worlds. It will report gate trigger counts on the target distribution as a first-class metric, and will automatically flag if the gate triggers zero times on positive-class examples (as in the paper’s failure). This prevents silent pipeline degradation where the gate never fires when needed.
Improvement 2: Implement dissociable failure decomposition for representation-to-behavior pipelines.
The improved system will separately track and report three independent obstacles: (a) whether the behaviorally-relevant readout exists, (b) whether representations align across evidence types, and (c) whether the answer is detectable and usable by the readout. It will not conflate these into a single “representation–behavior gap.” This allows targeted debugging: if localization succeeds but release fails, the system will isolate whether the bottleneck is the linear parametrization, the gate, or the readout, rather than attributing failure to “wrong layer” or “wrong direction.”
Improvement 3: Add dose–response plateau detection for linear release interventions.
The improved AI system will automatically test multiple intervention strengths (doses) and check whether the response plateaus below a predefined release threshold. If the plateau is detected (as in the paper’s intercept falling from 0.382 to 0.264 but never reaching 0.08), the system will flag that the linear release family is insufficient and will recommend switching to a non-linear or multi-site intervention, rather than continuing to tune the same linear direction.
Improvement 4: Require preregistered threshold and branch validation with hashed decision trees.
The improved system will enforce that all thresholds, claim templates, and decision-tree branches are hashed and archived before any data is collected, as done in the paper. This prevents post-hoc threshold tuning and ensures that any reported success or failure is statistically valid. The system will also automatically compute confidence intervals for release rates and trigger rates, and will flag degenerate statistical tests (e.g., McNemar with zero off-diagonal counts) as invalid.
Improvement 5: Incorporate target-distribution trigger counts into model selection and early stopping.
The improved system will use gate trigger counts on the target distribution (not just calibration OOD) as a primary metric during training or fine-tuning. If the gate triggers on zero positive-class examples (as in the paper’s 0/2,400), the system will halt and re-train the gate with positive-class examples included in calibration, preventing the “learned ‘looks like calibration OOD’ rather than ‘needs intervention’” failure.
Improvement 6: Add per-instance adaptivity bounds reporting for steering interventions.
The improved system will report the maximum and minimum per-instance adaptation effect (e.g., ±0.03 in the paper) alongside the fixed-direction effect. If the adaptive gain is negligible (as in the paper’s +0.0268 and −0.0199), the system will conclude that per-instance adaptivity does not overcome the linear parametrization bottleneck and will suggest exploring non-linear release families (e.g., learned rotations or multi-layer interventions).
What the improved AI system can do:
-
Reliably detect when a gate is silently failing on the target distribution, preventing wasted compute and false confidence.
-
Distinguish between localization success and release failure, enabling precise intervention design.
-
Automatically identify when linear steering is fundamentally insufficient, saving time by avoiding exhaustive tuning of linear directions.
-
Produce statistically valid, preregistered results with confidence intervals and trigger counts, reducing reproducibility failures.
-
Adaptively switch intervention strategies (e.g., from linear to non-linear) when dose–response plateaus indicate a parametrization bottleneck.
Abstract
A growing body of work reports that language models represent task-relevant latent structure that they fail to use. Whether such structure, once located, can be converted into behavior is a separate question that is rarely tested end to end. We submit the complete pipeline -- detect, localize, and release -- to a fully preregistered stress test on a 25.7M transformer trained on causal-evidence discrimination, where a known suppression phenomenon (latent causal structure present but behaviorally unused) has previously been documented. Every threshold, claim template, and decision-tree branch was hashed and archived before any corresponding data existed. Three findings. (i) Localization succeeds: interventions at observation-evidence channels of mid layers restore target behavior on otherwise-suppressed worlds (paired release advantages 0.563 and 0.854, 97.5% CIs excluding zero; best-site release rate 0.889). (ii) Gating fails out of distribution: a detector calibrated to trigger on zero out-of-distribution calibration worlds triggers on 6.9-7.3% of held-out in-distribution generations and on zero of the 2,400 held-out generations that actually need it -- a complete inversion that silently reduces the gated pipeline to its base model. (iii) Linear release is capped: removing the gate and injecting a per-instance linear direction unconditionally yields a monotone dose-response that plateaus far below the preregistered release margin (intercept 0.382 to 0.311 to 0.264 vs. threshold 0.08); per-instance adaptivity adds less than plus or minus 0.03. The failure is doubly located: the detector is OOD-inverted, and the entire family of linear release directions at this site and resolution is bounded away from sufficiency. The two failures are dissociable, and neither overturns localization. Every number traces to a hashed artifact in the released audit chain.
Sources
- Perfect Detection, Failed Control: The Geometry of Knowing vs. Steering in Language Models
- Steering Llama 2 via Contrastive Activation Addition
- Steering Language Models With Activation Engineering
- Evidence-Type Competition: When Can Interventional Data Teach Language Models Causal Direction?
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering