Causal Structure is Inducible but Functionally Decoupled: The Routing/Readout Boundary of a Typed Mechanism Library
Xining Xun
Tsingjiao Information Science (Beijing) Company Limited
cs.CL
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: 17 pages, 9 figures, 9 tables
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 100/100
The gist: Causal Structure Is Inducible but Functionally Decoupled: The Routing/Readout Boundary of a Typed Mechanism Library Authors: Xining Xun (Tsingjiao Information Science (Beijing) Co., Ltd.) arXiv:
Terminology
Summary
Causal Structure Is Inducible but Functionally Decoupled: The Routing/Readout Boundary of a Typed Mechanism Library
arXiv: 2608.11767v1 [cs.CL], 12 Aug 2026
The paper investigates whether exposing discrete internal structure in a language model — slots, neurons, circuits — and modifying it should show up in the model's output, an assumption the authors call the circuit-board intuition.
The authors built a controlled platform to test this assumption end to end, and the answer is no — in a precise, measured, and replicated sense.
The central finding is a routing/readout boundary (H-α): slot×type structure induced by type-level supervision organizes routing but remains functionally decoupled from answer readout. The paper reports that training decoupled a structure that was built to participate
— the architecture gave the explicit structure every condition to sit on the readout pathway, yet the optimizer separated addressing from computation.
The platform consists of:
-
A causal-world testbed with exact interventional ground truth, where every world is a small structural causal model and every query carries a known evidence type
-
A typed mechanism library (MM) on a standard transformer backbone: N discrete slots (N=600 at 125M, N=200 at 22.6M), statically partitioned over evidence types (6 types at 125M: identity/child/relation/sign/confidence/block-reserved; 3 types at 22.6M), with a per-type floor (βfloor = 0.3)
-
A gating head trained by a small auxiliary loss (λg = 0.1 for type-supervised mode) encouraging type-consistent routing
-
An EditSession interface with four structural operators (flip sign, add edge, remove edge, swap edge) plus a bounded parameter operator (param edit, ≤50 steps, row-masked lr 10−3)
-
A 116.66M-parameter causal-LM backbone (total 126.21M with library and heads at mm125), trained on online-sampled causal worlds with standard LM loss plus answer head, gating, and load-balancing terms
Three gate modes define the arm structure: type (supervised by type-classification auxiliary loss), emergent (λg = 0, free routing), and blocks (λg = 0, block-structured routing prior without type labels). The two unsupervised arms serve as the null model for the organization claim.
At 22.6M: Type-supervised routing produces slot×type organization replicating across three seeds: per-seed z = 4.47/5.99/7.02 (each p ≤ 5×10−4; Stouffer z = 10.1), debiased MI excess 0.072–0.124 nats (mean 0.0998). Both unsupervised arms sit on the null (emergent z=+0.21, blocks z=+0.14). The verdict is PASS at the moderate grade — the strong-grade mean-excess bar (0.10) was missed by 0.0002 and recorded as such. A control condition with semantically arbitrary labels was not learnable at all (archived control F1, two protocols, four seeds).
At 125M — the moving null: All three type seeds pass the raw amplitude gate (z = 6.09/7.75/5.62, excess = +0.141/+0.164/+0.134 nats), but the pipeline-validation gate did not pass cleanly: the blocks arm measured z=+2.38 (excess +0.050, p=0.012), above the z < 2 null band. The diagnostic finding: the unsupervised null is not scale-free. At 22.6M both unsupervised arms sit exactly on the null; at 125M the blocks arm carries a small, stable slot×type alignment. A fresh blocks seed settles the effect at arm-level (blocks s1: z=+3.89, excess +0.065 nats), not seed-level fluctuation. Because scale and type-count changed together (N 200→600, 3→6 types), the baseline shift's cause is confounded and reported as such.
Attribution revision: The binary pipeline check (unsupervised z < 2) assumes the unsupervised null is zero. When that assumption failed, the protocol revised the estimand, not the bar. The paired-permutation attribution test applies the same label permutation to both arms synchronously and normalizes the arm difference: z∆ = (∆ − null mean)/null sd over 2,000 paired permutations. On archived data (320 queries), five of six contrasts clear z∆ ≥ 3; the weakest (s2 vs blocks) reaches 2.68 (p=0.0035). The intersection gate was not met
; the edge clause opened exactly one remedy — a powered replication on fresh seeds with fresh worlds.
Replication (A20-R): Three fresh type seeds, two fresh unsupervised controls, 1,280 fresh queries, fresh world set. The gate passed on all nine criterion cells: amplitude z = 15.15/13.20/13.77; MI excess = +0.1275/+0.1025/+0.1202 nats; z∆ vs emergent s1 = +7.87/+5.92/+7.29; z∆ vs blocks s1 = +7.26/+5.02/+6.64. Replication control arms: emergent s1 z=7.58/excess +0.0471; blocks s1 z=10.88/excess +0.0573. A robustness-disclosure member (blocks s2, z=12.05, excess +0.0810) showed the unsupervised null strengthens further; the s4 cell landed below 3 (z∆ = +2.21, p=0.0145) and was reported as such per the anti-rescue clause.
Three findings scope every editability claim:
-
Applying all 150/150 structural flips changes the answer readout by ≤ 2×10−4 with exactly zero collateral — structural codes are not on the readout pathway
-
ROME-style row fine-tuning is plastic at the edited point (0.855) but does not generalize to held-out variants (0.124)
-
The tgC generalization probe fails its bar at both scales (0.124 at 22.6M; 0.217, three-seed mean, at 125M; frozen bar 0.30)
The paper states: We therefore describe the library as a typed routing index with exact state-level edit semantics, and we explicitly do not claim behavioral editability of the world model.
The decoupling is not an architectural triviality: the architecture gave the explicit structure every condition to participate in readout, and training decoupled it anyway.
LM-quality gap vs the parameter-matched monolith: 0.0058/0.0082/0.0071 nats across the three type seeds (criterion ≤ 0.02), and the 50-item Simpson-paradox behavioral battery passes 50/50 on both arms (criterion ≥ 95%). The typed library costs nothing measurable in base capability.
The library state is bit-exactly revertible as state. The preregistered R1–R5 audit passes on all three seeds with zero failures: 250/250 single-edit restores; 1,000/1,000 depth-20 stacked reverts; 1,250/1,250 backbone zero-touch comparisons; zero-contamination on every failed edit (147–150 sampled instances per seed). Rollback completeness holds by construction (the audit buffers live in the model's state dict, so no snapshot can forget revertible state).
Paired flip edits (150 queries/seed, three seeds) at 125M: answer-readout change ∆ŷ ≤ 3.4×10−6 (criterion ≤ 10−3, met with ≥ 2.5 orders of magnitude margin) and off-path collateral exactly 0.0 on all three seeds. The F6 pathway boundary — structural slot codes do not feed the answer readout — replicates at 5.6× scale without exception. The bounded parameter operator remains plastic at the training point (gain 0.884/0.935/0.917) but does not teach the readout to consume structural codes: tgC mean 0.217 < 0.30.
The paper reports that the unsupervised null itself moves with scale — a moving null — which scopes how organization claims may be inherited across scales.
A structural-prior control sits on the null at 22.6M but carries weak, stable organization at 125M (z=2.38). Architecture comparisons that reuse a null calibrated at one scale may be confounded at another.
The paper releases the paired-permutation attribution protocol as a reusable template for organization claims whose unsupervised null may be nonzero.
The paper proposes a hypothesis (explicitly labeled as speculation) for why decoupling exists: "Discrete slot codes are high-precision, low-bandwidth objects, well suited as addresses; coupling them directly to the readout would tie a sparsely-updated, discrete pathway to every output position, plausibly injecting gradient noise into the dense computation. A transformer trained by gradient descent may therefore find it favorable to keep discrete structure on the routing side while distributed weights carry the answer — a separation of addressing from computation that acts as an implicit regularizer, discovered by the optimizer rather than imposed by us."
-
Synthetic causal worlds; two scales within one architecture family
-
Three-seed arm budgets (six after replication)
-
The 22.6M organization effect is moderate-grade (three-seed mean excess 0.0998 against the frozen strong-tier bar of 0.10 — a 0.0002 miss recorded per the frozen tiering, not relabeled)
-
The baseline-emergence confound (§4.1) is reported, not resolved
-
Persistence under continued training (A21) remains untested — the criterion is frozen and was unfrozen for execution on 2026-08-06, but no A21 data exist yet
Every headline claim is paired with a preregistered, machine-checkable criterion archived (with md5-chained scripts) before the data it governs. Permutation tests (2,000 permutations, archived seeds) drive all mutual-information claims; editing claims are pass/fail with zero statistical degrees of freedom. A pipeline-validation gate accompanies every MI-based criterion: unsupervised control arms must sit within the null band (z < 2). The protocol specifies in advance what happens when a criterion fails: suspension of the affected claim, a frozen diagnostic battery, and — where warranted — a revised estimand whose thresholds are frozen before the data it governs. This path was executed once (the 125M pipeline-gate failure behind §4.1's branching). Anti-rescue discipline: no re-judging archived values, no threshold edits after data, no seed top-ups, no second replications.
Improvements for AI systems
Improvements to AI Systems:
-
Add a routing/readout decoupling layer to transformer architectures. Explicitly separate discrete slot-based addressing (routing) from continuous weight-based computation (readout) by adding a gating head trained with a small auxiliary loss (λg ≈ 0.1). This yields typed organization with zero measurable loss in base capability (gap ≤ 0.0082 nats vs. monolith) and bit-exact state reversibility.
-
Implement a typed mechanism library with static slot partitioning. Allocate N discrete slots (e.g., 600 at 125M scale) partitioned over semantic types (identity, child, relation, sign, confidence, block-reserved), each with a per-type floor (βfloor = 0.3). This induces slot×type structure via type-level supervision, achieving debiased mutual information excess of +0.10–0.16 nats with z-scores > 6.
-
Add an EditSession interface with structural operators and bounded parameter edits. Support flip sign, add edge, remove edge, swap edge, plus a row-masked parameter operator (≤50 steps, lr 10−3). This enables exact state-level edit semantics with answer-readout change ≤ 3.4×10−6 and zero off-path collateral, while keeping the model bit-exactly revertible (audit passes 250/250 single-edit restores, 1,000/1,000 depth-20 stacked reverts).
-
Integrate a paired-permutation attribution protocol for organization claims. Replace binary pipeline checks (unsupervised z < 2) with paired-permutation tests (2,000 permutations, archived seeds) that normalize arm differences. This handles moving nulls—where unsupervised baselines gain weak organization at larger scales (z = 2.38–12.05)—and prevents false attribution of induced vs. emergent structure.
-
Add a pipeline-validation gate with anti-rescue discipline. Pre-register machine-checkable criteria with md5-chained scripts before data collection. If a gate fails (e.g., unsupervised null exceeds z < 2), automatically suspend the claim, run a frozen diagnostic battery, and—if warranted—execute a powered replication with fresh seeds and fresh worlds, with thresholds frozen in advance. No re-judging, no seed top-ups, no second replications.
-
Include a generalization probe (tgC) for structural edits. After any structural edit, test whether the model generalizes to held-out variants. Current results show tgC fails (0.124–0.217 vs. 0.30 bar), so the improved system should explicitly report this and avoid claiming behavioral editability of the world model—only state-level edit semantics.
-
Add a Simpson-paradox behavioral battery (50 items). Validate that adding the typed library doesn't degrade reasoning under confounded evidence. The improved system passes 50/50 items, confirming the structure is free in capability cost.
What the Improved AI System Can Do:
-
Perform exact, reversible structural edits (flip, add, remove, swap edges) on its internal causal-world representation with zero behavioral side effects (∆ŷ ≤ 3.4×10−6) and full rollback via state-dict buffering.
-
Maintain typed, interpretable routing across 6 evidence types with high statistical confidence (z > 6), while keeping base LM quality within 0.0082 nats of a parameter-matched monolith.
-
Detect and correct for scale-dependent baselines in organization claims, avoiding false positives when comparing across model sizes (22.6M vs. 125M).
-
Guarantee auditability: every edit is bit-exactly revertible, every claim is tied to a pre-registered, machine-checkable criterion, and no post-hoc threshold adjustments are possible.
-
Avoid overclaiming editability: the system explicitly distinguishes state-level edit semantics from behavioral world-model editing, preventing misuse in safety-critical applications.
-
Scale robustly: the routing/readout boundary replicates at 5.6× scale, and the moving-null protocol prevents inherited confounds when scaling architecture comparisons.
Abstract
When a language model answers an interventional question, the computation it must perform depends on the type of evidence the query requires. We report a decoupling in how a transformer organizes causal knowledge: slot-by-type structure induced by type-level supervision organizes routing, yet remains functionally decoupled from answer readout. We establish this with a typed mechanism library -- discrete mechanism slots partitioned by evidence type, auditable at the state level -- on a causal-world benchmark with exact interventional ground truth, under a frozen protocol, at two scales (22.6M and 125M). Four preregistered findings. (i) Origin. Slot-by-type organization is induced by type-level supervision: absent in architecturally identical unsupervised controls, not buyable by content-free gating labels, and statistically attributable to the supervision signal, replicating at 125M under a powered preregistered protocol (all nine cells passed). (ii) Boundary. The induced structure is a typed routing index with a sharp routing/readout boundary: slot codes scaffold routing but do not drive answer readout (3.4 times10-6, zero collateral, three seeds, stable across a 5.6x scale window) -- we therefore make no behavioral-editability claim. (iii) Cost. The structure is free: LM quality matches a parameter-matched monolith within 0.0082 nats. (iv) Trust. The library state is exactly local under edit and bit-exactly revertible -- 250 single-edit and 1,000 stacked reverts per seed, zero failures. We further find that the unsupervised null itself moves with scale, so comparisons reusing a null calibrated at one scale may be confounded at another. Every claim is tied to a preregistered, machine-checkable criterion archived before the data it governs; the full audit trail, including one criterion we failed and how the frozen protocol handled it, is released as an appendix.
Sources
- The Scientific Method in the Science of Machine Learning
- Neural Turing Machines
- Causal Reasoning and Large Language Models: Opening a New Frontier for Causality
- Evidence-Type Competition: When Can Interventional Data Teach Language Models Causal Direction?
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering