Reading Between the Dots: Decoding Hidden Computation across Filler Tokens
summary
The gist
Frontier Large Language Models can perform multi-step reasoning over content-free filler tokens, and this hidden computation is at least partially accessible to interpretability tools when it is
In short
Researchers found that large language models perform hidden reasoning over filler tokens (like dots) to improve performance. They developed an unsupervised decoding pipeline to recover these intermediate values from the model's residual stream. This shows that behavioral monitoring is not the only way to audit models, as hidden computation is partially accessible through this method.
Key concepts
- Filler Tokens
- These are content-free tokens, such as dots or counting sequences, placed between a question and an answer slot in a prompt. The study investigates whether these tokens enable models to perform complex reasoning steps that lead to better answers.
- Unsupervised Decoding Pipeline
- This is a four-stage process designed to extract hidden intermediate values from the model's internal states. It involves extracting residual streams at every layer and filler position, applying logit lenses, aggregating scores across positions, and using another LLM judge to decode the reasoning.
- Residual Stream
- This refers to the leftover information or 'noise' in a transformer model after it has processed an input. The paper suggests that this residual stream contains structured information about the model's hidden computation, which can be decoded to find intermediate values.
- Logit Lens
- A tool used during the decoding process to obtain a probability distribution over vocabulary at specific points within the model's layers and filler tokens. It helps researchers pinpoint which tokens are most likely given the model's internal state, aiding in value extraction.
Terminology used across episodes
This episode discusses
- Reading Between the Dots: Decoding Hidden Computation across Filler Tokens · Paper Radio
- Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop Queries
- Reasoning Models Don't Always Say What They Think
- DeepSeek-V3 Technical Report
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
- How to use and interpret activation patching
- Kimi K2: Open Agentic Intelligence
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
- Measuring Faithfulness in Chain-of-Thought Reasoning
The paper
Reading Between the Dots: Decoding Hidden Computation across Filler Tokens · Read on arXiv
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Reading Between the Dots".
Jane: Frontier Large Language Models can perform multi-step reasoning over content-free filler tokens,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: We’ve established that the authors of "Reading Between the Dots: Decoding Hidden Computation across Filler Tokens" are looking at how models use filler tokens to perform complex reasoning, and they aren't just observing this; they’re trying to decode what's happening under the hood.
Jane: That's right. The title itself points directly to the core idea: we need a way to read between the dots, meaning we can extract meaningful information from those filler tokens that aren't visible in a standard output format.
Lu: What really catches my eye is their focus on four specific task families—fact retrieval, parallel numeric composition, string manipulation, and in-context computation—showing that this isn't just an isolated quirk but a general behavior across different kinds of problems.
Meng: So it’s not limited to one type of math problem; it shows up in a variety of cognitive tasks, which makes the finding much more broadly applicable than I initially thought when I was thinking about practical deployment.
Lalam: It suggests that we should start thinking less about just what the final answer is and more about the structured process happening within those filler tokens themselves, which opens up new avenues for system introspection.
The paper's summary: Tom: The main summary of "Reading Between the Dots" shows that these frontier LLMs use filler tokens to perform multi-step reasoning, and this computation is structured in a way that attention routes the question through the filler to the answer.
Jane: That’s a key mechanism they highlight: the attention patterns aren't just random; they form a relay where the question flows through those dots, which then leads directly to the final answer slot.
Lu: They also point out that when you look at logit-lens readouts, you can see retrieved facts appearing early in those filler regions and their composition building up in later layers right before the answer is generated.
Meng: That temporal progression across layers sounds like a very organized way for the model to handle sequential steps, which is something we need to understand if we want to build reliable reasoning engines.
Lalam: It’s encouraging because they show that this isn't just accidental correlation; the researchers have an unsupervised decoding pipeline that can recover these intermediate values with high accuracy, up to ninety-five percent.
The paper's improvements: Tom: The improvement they introduce is this unsupervised decoding pipeline, which takes only hidden states as input and recovers intermediate values with eighty to ninety-five percent accuracy across both models and all four task families.
Jane: So, the method involves extracting the residual stream at every layer and position, applying a logit lens to get a probability distribution at each cell, and then aggregating those scores to find the most important tokens.
Lu: Their mechanistic evidence is pretty strong too; they show that attention analysis confirms this relay structure, while KV-cache transplants provide causal proof that the content in those filler positions actually drives the final answer in specific examples.
Meng: Causal reliance is important for me because it moves beyond mere correlation; it proves that if we manipulate that filler content, we can reliably expect a predictable change in the outcome of the model.
Lalam: This pipeline is incredibly powerful because it’s unsupervised; they don't need us to give them ground-truth labels to teach the system how to read these hidden computations, which makes it very general.
Conclusion: Tom: So, we’re wrapping up with a look at what this means for the field. Essentially, "Reading Between the Dots: Decoding Hidden Computation across Filler Tokens" shows that hidden computation over filler tokens is a general phenomenon and that an unsupervised pipeline can decode those intermediate values with eighty to ninety-five percent accuracy across both models and all four tasks.
Jane: It really solidifies the idea that behavioral monitoring isn't the only way to audit these systems; there’s a deeper, hidden layer of reasoning we can access through this residual stream analysis.
Lu: The implication is significant because it shows that for tasks with clean intermediate values—like numbers or entity names—we have a concrete method to interpret what the model was actually reasoning about internally.
Meng: For us in engineering, this means we can better diagnose where a computation fails; for instance, if retrieval works but composition doesn't, we know exactly where the breakdown occurred in the hidden stream.
Lalam: This work suggests that monitorability is a property of the model’s full computational trace rather than just its surface output, which is a huge step toward building more trustworthy and explainable AI.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language