Where You Measure Decides What You Measure: Position Selection in Ablation-Based SAE Evaluation

arXiv:2608.13337 · cs.LG · Submitted 2026-08-13 · Read on arXiv

Valentin Noël

Devoteam

cs.LG

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: 19 pages, 3 figures. Code and data: https://github.com/vcnoel/sae-artifact

Code: https://github.com/vcnoel/sae-artifact

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 100/100

The gist: Author: Valentin Noël (Devoteam) arXiv: 2608.13337v1 [cs.LG], 13 Aug 2026 --- The paper investigates a hidden methodological flaw in how sparse autoencoder (SAE) latents are causally evaluated.

Terminology

Summary

Author: Valentin Noël (Devoteam)

arXiv: 2608.13337v1 [cs.LG], 13 Aug 2026


The paper investigates a hidden methodological flaw in how sparse autoencoder (SAE) latents are causally evaluated. The abstract states: "Sparse autoencoders are meant to name the things a language model computes, and the usual way to check that a latent matters is to switch it off and see what changes. But a latent fires at many tokens, and the effect has to be measured at one of them. The convention is to measure where the latent fires hardest. That choice is almost never reported, and it is not made by the experimenter: it is made by the dictionary under evaluation. Change the dictionary and the measurement moves to a different token."

The key insight is that "the token where a latent fires hardest is not a fact about the model. It is a fact about the dictionary, computed from the dictionary's own activations. Fit a second autoencoder on the same data and it will place the same latent's maximum somewhere else."

Using two Gemma Scope autoencoders released by Google for the same base model, the author matched latents by decoder similarity. Even among pairs the two dictionaries encode almost identically, they pick different tokens for a large share of them. Across 15 pairs of released Gemma Scope dictionaries, no pair agrees on the measurement token for more than 48% of its matched latents, and even restricting to the latents a pair encodes almost identically the figure only reaches 60.2%.

The paper reports: on this corpus a released dictionary's latents are live at a median of 103 positions each, ranging from 31 to 367 across the six dictionaries. Agreement is above chance (13.9% observed vs 3.5% under a shuffle null), but the arms agree about four times as often as chance — far below what a comparison assumes.

To separate the convention from the dictionaries, the author trained six autoencoders from one shared initialisation, differing only in fitting choices (decoder free, soft-frozen at τ=0.80, soft-frozen at τ=0.90, learning rate 10× lower, sparsity k=41 instead of 82, or reshuffled corpus order). This ensures latent i denotes the same thing in all six and any disagreement between them is attributable rather than confounded.

The key result: "Most of the variance such a comparison reads as these dictionaries disagree about this latent turns out to be the position instead: it falls from 7.6% and 11.9% of variance to near zero once every dictionary is measured at the same token."

Specifically, Table 1 shows:

  • Gemma-2-2B: latent×arm variance drops from 7.6% (per-arm positions) to 0.0% (shared positions); Eρ2 rises from 0.738 to 0.869

  • Gemma-3-1B: latent×arm variance drops from 11.9% to 2.4%; Eρ2 rises from 0.634 to 0.767

The paired gain in Eρ2 is +0.130 (95% CI [+0.074, +0.208]) on Gemma-2-2B and +0.138 ([+0.024, +0.290]) on Gemma-3-1B.

More evaluation data does not rescue it. Across a sixteenfold range of corpus sizes the dictionaries agree less about where to measure, not more, so the problem grows with scale.

Position agreement falls monotonically: 18.1% → 13.9% → 10.0% on Gemma-2-2B and 13.6% → 11.3% → 9.4% on Gemma-3-1B across 96, 384, and 1536 sequences. The paired gain from controlling position grows: +0.082, +0.130 and +0.186 on Gemma-2-2B and +0.046, +0.138 and +0.362 on Gemma-3-1B.

The within-cell term of Table 1 is position within latent, and it is 67.4% of variance, larger than latent, arm, and their interaction combined. This means the token a latent is measured at moves the answer more than which dictionary produced the latent.

However, regressing this term on relative position, activation magnitude and within-cell activation rank together explains R2 = 0.005 of it, so no obvious covariate substitutes for the choice.

Whether the readout is normalised by the size of the intervention... can flip the sign of a real comparison on identical data. Specifically: "asked whether a dictionary's rarest activation-frequency decile carries more causal mass than a frequency-matched decile of a tied-random dictionary, raw KL says worse (HL = 0.683, 95% CI [0.459, 0.961], p = 0.039) and KL per unit norm says better (HL = 1.687, [1.171, 2.498], p = 0.011), on the identical latents and the identical control, both intervals excluding one."

"The BOS residual carries an activation the dictionary does not reconstruct, cosine 0.44 between input and reconstruction at position 0 against ≈ 0.93 elsewhere, and its norm dominates any pooled statistic: including position 0 moves the explained variance of a released Gemma Scope dictionary from 0.863 to −3.5."

The paper states: "A practitioner reporting an ablation-based causal effect should:

  1. Report the position, and how many positions per latent were measured.

  2. Report effect per unit perturbation norm, not raw.

  3. Drop special-token positions before any pooled statistic, and assert the expected explained variance rather than trusting it.

  4. Report the range across fitting choices, not a single contrast."

Five audited papers report a single-latent causal effect with a magnitude readout; none report the position at which it was taken. The audited papers are Bricken et al. 2023, Marks et al. 2024, Templeton et al. 2024, Gao et al. 2024, and Cho et al. 2026. The audit records what is reported, not what was done: an unreported control cannot be checked or reproduced, whoever omitted it.

The author acknowledges: Two models, one family. Both base models are Gemma; whether the convention behaves the same in another architecture is untested. The controlled design measures only latents firing in every arm at a common position, 55% and 35% of those evaluated, and these latents have systematically lower measurement error than the ones it excludes. The results concern metrics that zero a latent's contribution and read a magnitude; they do not transfer directly to interchange-based scores such as RAVEL.

Training scale is a limitation for the repair but not the premise: "The premise, that two dictionaries select different measurement positions for the same latent, needs no shared initialisation and no training of ours... The repair, the crossed decomposition and the Eρ2 gain, cannot be measured this way, and the reason is structural rather than a shortfall of ambition."

The paper concludes: "Because the convention selects that position from the dictionary's own activations, two dictionaries under comparison are measured at different tokens, agreeing on the top-activating position for 13.9% and 11.3% of latents, less as the corpus grows. Most of the latent×arm variance such a comparison attributes to the dictionaries belongs to that instead: on identical latents it falls from 7.6% to 0.0% and from 11.9% to 2.4% once position is fixed. None of this depends on dictionaries we trained: across 15 pairs of released Gemma Scope dictionaries, no pair agrees on the measurement token for more than 48% of its matched latents, and even restricting to the latents a pair encodes almost identically the figure only reaches 60.2%. The correction is one line of evaluation code, checkable in an afternoon."

The paper ends with an open question: "if a latent's effect varies this much across the tokens where it fires, the quantity to report may be a distribution over those tokens, an activation-weighted expectation, or a context-conditioned estimand rather than a scalar, decidable empirically, by whichever is more stable across dictionaries at equal cost on the crossed design used here."

Improvements for AI systems

Improvements to AI systems:

  1. Position-aware causal evaluation module: Add a component that automatically identifies and reports the measurement position for every SAE latent ablation, rather than defaulting to the max-activation token. The system tracks how many positions each latent was measured at and flags when position choice could confound results.

  2. Cross-dictionary position agreement checker: Implement a validation step that, when comparing two SAEs, computes the fraction of matched latents where both dictionaries agree on the measurement token. The system warns the user if agreement falls below a configurable threshold (e.g., 60%) and suggests reporting the range of effects across positions instead of a single scalar.

  3. Position-controlled variance decomposition: Add a diagnostic that decomposes variance in causal effect estimates into components attributable to latent identity, dictionary choice, and measurement position. The system automatically reports the latent×arm interaction term both with per-arm positions and with shared positions, making it obvious when position choice dominates.

  4. Special-token exclusion filter: Build in automatic detection and removal of special-token positions (e.g., BOS) before computing any pooled statistics. The system checks reconstruction cosine similarity per position and flags positions where the dictionary fails to reconstruct, preventing artifacts like the explained variance dropping from 0.863 to −3.5.

  5. Normalization-aware reporting: Add a toggle that reports causal effects both as raw KL divergence and as KL per unit intervention norm. The system surfaces cases where these two metrics disagree in sign, alerting the user to a potential reporting-choice confound.

  6. Fitting-choice sensitivity range: Implement a multi-arm evaluation that trains or loads several SAEs with different fitting choices (learning rate, sparsity, corpus order) and reports the range of causal effects across them, rather than a single contrast. The system flags when the range spans zero or flips sign.

  7. Corpus-size scaling monitor: Add a check that tracks position agreement as evaluation corpus size grows. The system warns if agreement decreases with more data, indicating that the problem worsens with scale and that the user should not expect more data to fix it.

  8. Distribution-over-positions estimator: Offer an alternative output format that reports the causal effect as a distribution over the tokens where the latent fires, along with an activation-weighted expectation and a context-conditioned estimate. The system can compare which of these is more stable across dictionaries at equal computational cost.

What the improved AI system can do:

  • Produce causal effect estimates for SAE latents that are robust to the choice of measurement position, with confidence intervals that account for position variability.

  • Automatically flag when two dictionaries under comparison are being measured at different tokens, and quantify how much of the observed disagreement is attributable to position rather than dictionary content.

  • Generate evaluation reports that include all four protocol elements (position reporting, per-unit-norm effects, special-token exclusion, fitting-choice ranges) without requiring the user to manually implement them.

  • Detect when more evaluation data would worsen position agreement, saving compute by avoiding futile scaling.

  • Provide a principled way to choose between scalar, distributional, or context-conditioned effect estimates based on which is more stable across dictionaries, enabling more reliable SAE comparisons in interpretability research.

Abstract

Sparse autoencoders are meant to name the things a language model computes, and the usual way to check that a latent matters is to switch it off and see what changes. But a latent fires at many tokens, and the effect has to be measured at one of them. The convention is to measure where the latent fires hardest. That choice is almost never reported, and it is not made by the experimenter: it is made by the dictionary under evaluation. Change the dictionary and the measurement moves to a different token. We show this is not a detail. Take two sparse autoencoders released by Google for the same model and match their latents by decoder similarity: even among the pairs the two dictionaries encode almost identically, they pick different tokens for a large share of them. Two dictionaries compared under the usual protocol are therefore very often compared at different places. To separate the convention from the dictionaries we train six autoencoders from one initialisation, differing only in fitting choices, so that a latent means the same thing in each. Most of the variance such a comparison reads as "these dictionaries disagree about this latent" turns out to be the position instead: it falls from 7.6% and 11.9% of variance to near zero once every dictionary is measured at the same token. More evaluation data does not rescue it. Across a sixteenfold range of corpus sizes the dictionaries agree less about where to measure, not more, so the problem grows with scale. The correction is one line of evaluation code. We give the protocol an ablation-based causal number must report to be comparable across papers, and an audit of five published papers against it. In short: a causal number reported without its position describes the token it was taken at as much as the latent it was taken from.

Sources

Related papers