Reading Between the Dots: Decoding Hidden Computation across Filler Tokens
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Reading Between the Dots".
Jane: Frontier Large Language Models can perform multi-step reasoning over content-free filler tokens,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: We’ve established that the authors of "Reading Between the Dots: Decoding Hidden Computation across Filler Tokens" are looking at how models use filler tokens to perform complex reasoning, and they aren't just observing this; they’re trying to decode what's happening under the hood.
Jane: That's right. The title itself points directly to the core idea: we need a way to read between the dots, meaning we can extract meaningful information from those filler tokens that aren't visible in a standard output format.
Lu: What really catches my eye is their focus on four specific task families—fact retrieval, parallel numeric composition, string manipulation, and in-context computation—showing that this isn't just an isolated quirk but a general behavior across different kinds of problems.
Meng: So it’s not limited to one type of math problem; it shows up in a variety of cognitive tasks, which makes the finding much more broadly applicable than I initially thought when I was thinking about practical deployment.
Lalam: It suggests that we should start thinking less about just what the final answer is and more about the structured process happening within those filler tokens themselves, which opens up new avenues for system introspection.
The paper's summary: Tom: The main summary of "Reading Between the Dots" shows that these frontier LLMs use filler tokens to perform multi-step reasoning, and this computation is structured in a way that attention routes the question through the filler to the answer.
Jane: That’s a key mechanism they highlight: the attention patterns aren't just random; they form a relay where the question flows through those dots, which then leads directly to the final answer slot.
Lu: They also point out that when you look at logit-lens readouts, you can see retrieved facts appearing early in those filler regions and their composition building up in later layers right before the answer is generated.
Meng: That temporal progression across layers sounds like a very organized way for the model to handle sequential steps, which is something we need to understand if we want to build reliable reasoning engines.
Lalam: It’s encouraging because they show that this isn't just accidental correlation; the researchers have an unsupervised decoding pipeline that can recover these intermediate values with high accuracy, up to ninety-five percent.
The paper's improvements: Tom: The improvement they introduce is this unsupervised decoding pipeline, which takes only hidden states as input and recovers intermediate values with eighty to ninety-five percent accuracy across both models and all four task families.
Jane: So, the method involves extracting the residual stream at every layer and position, applying a logit lens to get a probability distribution at each cell, and then aggregating those scores to find the most important tokens.
Lu: Their mechanistic evidence is pretty strong too; they show that attention analysis confirms this relay structure, while KV-cache transplants provide causal proof that the content in those filler positions actually drives the final answer in specific examples.
Meng: Causal reliance is important for me because it moves beyond mere correlation; it proves that if we manipulate that filler content, we can reliably expect a predictable change in the outcome of the model.
Lalam: This pipeline is incredibly powerful because it’s unsupervised; they don't need us to give them ground-truth labels to teach the system how to read these hidden computations, which makes it very general.
Conclusion: Tom: So, we’re wrapping up with a look at what this means for the field. Essentially, "Reading Between the Dots: Decoding Hidden Computation across Filler Tokens" shows that hidden computation over filler tokens is a general phenomenon and that an unsupervised pipeline can decode those intermediate values with eighty to ninety-five percent accuracy across both models and all four tasks.
Jane: It really solidifies the idea that behavioral monitoring isn't the only way to audit these systems; there’s a deeper, hidden layer of reasoning we can access through this residual stream analysis.
Lu: The implication is significant because it shows that for tasks with clean intermediate values—like numbers or entity names—we have a concrete method to interpret what the model was actually reasoning about internally.
Meng: For us in engineering, this means we can better diagnose where a computation fails; for instance, if retrieval works but composition doesn't, we know exactly where the breakdown occurred in the hidden stream.
Lalam: This work suggests that monitorability is a property of the model’s full computational trace rather than just its surface output, which is a huge step toward building more trustworthy and explainable AI.
cs.CL, cs.AI, cs.LG
Submitted: 2026-07-03
Updated: 2026-10-01
Code: https://github.com/kaleybrauer/filler-token-reasoning
Importance score: 92/100
The gist: Frontier Large Language Models can perform multi-step reasoning over content-free filler tokens, and this hidden computation is at least partially accessible to interpretability tools when it is
Key concepts
- Filler Tokens
- These are content-free tokens, such as dots or counting sequences, placed between a question and an answer slot in a prompt. The study investigates whether these tokens enable models to perform complex reasoning steps that lead to better answers.
- Unsupervised Decoding Pipeline
- This is a four-stage process designed to extract hidden intermediate values from the model's internal states. It involves extracting residual streams at every layer and filler position, applying logit lenses, aggregating scores across positions, and using another LLM judge to decode the reasoning.
- Residual Stream
- This refers to the leftover information or 'noise' in a transformer model after it has processed an input. The paper suggests that this residual stream contains structured information about the model's hidden computation, which can be decoded to find intermediate values.
- Logit Lens
- A tool used during the decoding process to obtain a probability distribution over vocabulary at specific points within the model's layers and filler tokens. It helps researchers pinpoint which tokens are most likely given the model's internal state, aiding in value extraction.
Terminology
Summary
Frontier Large Language Models can perform multi-step reasoning over content-free filler tokens, and this hidden computation is at least partially accessible to interpretability tools when it is invisible to chain-of-thought monitoring. This work empirically demonstrates that models compute meaningful intermediate values across filler tokens in a structured way, which can be recovered from the residual stream using an unsupervised decoding pipeline, suggesting that behavioral monitoring is not the only auditing tool available.
How it works
The researchers study open-weights frontier models (DeepSeek V3 and Kimi K2) on four task families: 1-fact addition, parallel numeric composition, string manipulation, and in-context computation. They append content-free filler tokens (dots, counting sequences, alphabet sequences) between the question and the answer slot to observe performance uplift. The key finding is that filler tokens improve performance across all tasks; for instance, DeepSeek V3 climbs from a 54% baseline to ∼72% on 1-fact addition with 100–500 dots or counting tokens.
The paper introduces an unsupervised decoding pipeline designed to recover intermediate values from hidden states. This pipeline involves four stages:
-
Hidden-state extraction: Extracting the residual stream at every transformer layer and every filler token position for each example.
-
Residual logit lens per (layer, filler position): Applying the logit lens to obtain a probability distribution over the vocabulary at each cell and calculating a per-example residual by subtracting the cross-example mean. The top-T tokens are saved by residual score per example per (layer, position).
-
Aggregation across filler positions: Summing the residual scores across all (layer, position) settings for each example and ranking tokens globally by this aggregated score to keep the top 50.
-
LLM decode: Passing the top-50 list to an LLM judge (Claude Haiku or Claude Sonnet) under a neutral prompt asking what specific number(s) or concept(s) the model was reasoning about, with scores assigned based on whether the ground-truth intermediate appears in the judge’s top-K predictions.
Mechanistic Evidence for Encoding
The paper provides three lines of evidence showing how filler positions encode task-relevant information.
-
Attention analysis shows that when filler is present, the model
redistributes its attention to form a processing relay of question → filler → answer that replaces the answer’s direct line to the question.
-
Logit-lens readouts reveal what the filler region encodes:
across tasks and models, the intermediate values of the computation appear in the residual stream.
For 2-fact addition, A1 is encoded strongly early in the filler, A2 later, and their sum crystallizes in late layers immediately before the answer. -
KV-cache transplants establish causal reliance on this content: transplanting filler KV cache between matched examples
drives the donor’s answer up by tens to hundreds of ranks,
proving thatfiller-position content is causal for the final answer.
Decoding Results and Diagnostic Capabilities
The unsupervised pipeline achieves 80–95% accuracy in recovering intermediate values across both models and all four tasks. For 2-fact addition, the pipeline recovers both addends A1 and A2 in the top-2 numeric tokens in 35.2% of examples for DeepSeek V3, with judges raising this to 82.3% (Sonnet). The decoder is diagnostic about where computation fails; for instance, when retrieval succeeds but composition does not (the 2-fact failure mode), the operands occupy them more readily,
while the sum is absent from the residual stream.
Limitations and Future Directions
The pipeline targets tasks with discrete, nameable intermediate values (a number, an element name, a city),
recovering these with high accuracy. However, it cannot recover computation that lives in attention patterns or orthogonal directions to the vocabulary projection. The paper notes that for harder intermediates like strategies or goals (an intent to deceive
), the pipeline may not be accessible unless extended via trained probes or patchscope-style elicitation.
Furthermore, residualization is not uniformly beneficial; for counting filler, it can hurt the Haiku judge by up to 11.5 pp,
suggesting that raw logit-lens readouts may be preferable in certain cases. The paper concludes that hidden computation is at least partially accessible to interpretability on tasks with clean intermediates.
The gist: Hidden computation over filler tokens is a general phenomenon, and an unsupervised pipeline can decode the intermediate values with 80–95% accuracy across both models and all four tasks, suggesting monitorability is a property of the model’s full computational trace. The paper shows that filler positions causally carry task-relevant information, and this hidden computation is at least partially accessible to interpretability on tasks with clean intermediates.
Improvements for AI systems
Based on the scientific paper Reading Between the Dots: Decoding Hidden Computation across Filler Tokens,
here are specific, actionable improvements for AI systems and what those improved systems can achieve:
The core capability derived from this research is a method to extract hidden, multi-step reasoning from models that deliberately obfuscate their chain-of-thought (CoT) by using content-free filler tokens.
Here are the specific improvements and capabilities:
-
Implementation of the Unsupervised Decoding Pipeline for Hidden Computation Recovery
-
Mechanism: Implement a pipeline that takes only hidden states as input, applies a logit lens at every (layer, filler position), subtracts the cross-example mean to remove noise, aggregates residual scores across all positions and layers, and feeds the top 50 tokens to an LLM judge for interpretation.
-
Capability: This allows the system to
read
what a model is reasoning about in hidden states without requiring ground-truth labels or training data, achieving 80–95% accuracy on recovering intermediate values (numeric, string-valued, and computed). -
Specific Application: Enables the system to reliably decode multi-step arithmetic problems (e.g., 2-fact addition), complex logic puzzles, or mathematical systems of equations even when the model uses filler tokens instead of explicit CoT steps to solve them correctly.
-
Causal Reasoning via KV-Cache Transplants
-
Mechanism: Implement a causal intervention where the Key/Value (KV) cache content from a
donor
example (which has the correct answer) is transplanted into thetarget
example's forward pass, specifically at the filler token positions identified as encoding intermediate values. -
Capability: This verifies that the content in those filler positions is not just incidental noise but is causally responsible for deriving the final answer, leading to
full-swap
rates (up to 22% for certain tasks) when correctly localized. -
Specific Application: Allows developers to confirm that specific filler sequences are functionally equivalent to
hidden chain-of-thought
steps, validating the model's internal reasoning structure rather than just its surface output. -
Diagnostic Failure Mode Identification
-
Mechanism: Utilize the residual stream analysis to determine whether a model fails due to
missing retrieval
(e.g., failing to retrieve an intermediate value) versusmissing composition
(e.g., retrieving parts but failing to sum them). -
Capability: This provides a diagnostic tool for debugging and safety auditing, allowing researchers to pinpoint the exact stage of computation where the LLM reasoning broke down, moving beyond simple behavioral CoT monitoring.
-
Hybrid Auditing System (Behavioral + Interpretability)
-
Mechanism: Integrate the residual-stream decoding results with traditional Chain-of-Thought monitoring (behavioral oversight).
-
Capability: Creates a more robust safety layer where surface token monitoring identifies suspicious behavior, and residual stream analysis confirms whether that behavior stems from intentional obfuscation or genuine algorithmic failure.
-
Targeted Intermediate Recovery (Nameable Values)
-
Capability: The system can recover specific, discrete intermediate values like element names, city names, or specific numbers with high accuracy (80–95%) when the computation involves retrieving a single named entity.
Sources
- Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop Queries
- Reasoning Models Don't Always Say What They Think
- DeepSeek-V3 Technical Report
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
- How to use and interpret activation patching
- Kimi K2: Open Agentic Intelligence
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
- Measuring Faithfulness in Chain-of-Thought Reasoning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering