Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens
cs.CL, cs.AI, cs.LG
Submitted: 2026-09-01
Updated: 2026-09-01
Comments: 55 pages, 33 figures, 42 tables. Under review. Code: https://github.com/hematteo/sparse-readout-prism Dictionaries: https://huggingface.co/hematteo/sparse-readout-prism
Code: https://github.com/hematteo/sparse-readout-prism
Project page: https://neuralblog.github.io/logit-prisms
License: http://creativecommons.org/licenses/by/4.0/
The gist: A language model's prediction of its next token develops across layers, and lens methods track this process by decoding intermediate hidden states into tokens.
Terminology
Abstract
A language model's prediction of its next token develops across layers, and lens methods track this process by decoding intermediate hidden states into tokens. But a lens reading reflects both the hidden state and the readout (the unembedding matrix) used to decode it. Many lenses are fit on a corpus, and we show that two lenses differing only in their fitting corpus can report different tokens for the same hidden states. We call this dependence corpus conditionality. To examine readout structure independently of the fitting corpus, we introduce Sparse Readout Prism (SRP), which decomposes the readout using only its weights and expresses any token logit or logit difference as a sum of contributions from sparse readout features. This reveals readout features as a new unit of analysis for lens readings, exposing structure that token identities can obscure and enabling comparisons across tokens, contexts, layers, and lenses. Replacing the original readout with SRP's sparse approximation reconstructs 8.9-17.3 percentage points more of the tested logit differences than the strongest of six baselines built on geometric relations among readout rows. Ablating features shifts logit differences in proportion to their SRP contributions. Although token readings vary with the fitting corpus, the dominant readout feature remains stable. Because SRP uses no corpus in its construction, it provides a control independent of the fitting corpus for lens analyses.
Sources
- BatchTopK Sparse Autoencoders
- Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition
- Weight-sparse transformers have interpretable circuits
- LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs
- Ministral 3
- Disentangling Dense Embeddings with Sparse Autoencoders
- Beyond Activation Patterns: A Weight-Based Out-of-Context Explanation of Sparse Autoencoder Features
- The Llama 3 Herd of Models
- Output Embedding Centering for Stable LLM Pretraining
- Qwen3 Technical Report
- LogitLens4LLMs: Extending Logit Lens Analysis to Modern Large Language Models
- Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
- Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
- Do Multilingual LLMs Think In English?
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering