Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens".
Jane: The paper was written by Zhenyu Zhao, Sander Land, Daniel M. Bikel and Waseem Alshikh from Lila Sciences and Writer, Inc..
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: We are looking at a really interesting piece today called "Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens."
Jane: That title makes it sound like we are teaching these models to use text abbreviations when they think.
Lu: It is actually a much more sophisticated version of that idea, Jane.
Tom: The authors, including Zhenyu Zhao from Lila Sciences and several researchers from Writer, Inc., seem to be targeting something very specific.
Jane: They are looking at the way these models talk to themselves during a reasoning process.
Meng: I am curious if they are actually changing the model or just the way we read the output.
Lu: They are changing the vocabulary itself, which is a massive shift in how we approach efficiency.
Tom: Using "supertokens" as a concept implies they are bundling multiple small pieces into one big unit.
Jane: So instead of the model saying "Let," "us," "verify," "that," it just uses one single token for the whole phrase?
Meng: That sounds like it could save a lot of processing time if it works.
Lu: It is the "entropy-guided" part of the title that really catches my eye.
Tom: Right, because they aren't just picking random phrases to bundle together.
Jane: They are using math to find the phrases that are the most predictable.
Meng: If a phrase is predictable, it means it doesn't carry much new information, does it?
Lu: Exactly, and that is where the compression happens.
Tom: I wonder how they actually decide which phrases are worth turning into a supertoken.
Jane: That is exactly what the summary of their method covers.
Summary: Tom: We are continuing our look at "Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens," focusing on how they actually do it.
Jane: They realized that reasoning tokens aren't all the same.
Lu: They split them into two distinct groups: structural tokens and organic tokens.
Tom: I love that distinction, especially the idea of "organic" tokens being the actual math or logic.
Jane: Think of structural tokens as the scaffolding, like when a model says, "Wait, let me reconsider."
Meng: So the "organic" tokens are the actual meat of the problem, like the variable names or the numbers?
Lu: Precisely, and the structural ones are very low in entropy because they follow a set pattern.
Tom: Once the model starts a phrase like "Let us check," the next few words are almost guaranteed.
Jane: They use a process called BPE merges on the model's own reasoning traces to find these patterns.
Meng: Are they training the model from scratch to use these new tokens?
Lu: No, they are just doing a lightweight fine-tuning on the embedding layer and a few transformer layers.
Tom: It is a much more surgical approach than retraining a whole model.
Jane: It is like giving the model a new set of specialized pens that can write entire sentences in one stroke.
Meng: I am wondering if this makes the reasoning process harder for the model to follow.
Lu: It actually makes it more efficient because the model isn't wasting energy on the predictable parts.
Tom: They even show that these supertokens align with actual reasoning moves like backtracking or verification.
Jane: It is like the model is using signposts to tell itself where it is in the logic.
Meng: That sounds like it might lead to some impressive performance gains.
Improvements: Tom: We have been talking about the mechanics of "Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens," and now we need to look at the actual results.
Jane: The numbers they reported are quite impressive, especially the eight point one percent average reduction in reasoning tokens.
Lu: That might not sound huge, but in the world of massive reasoning traces, it is significant.
Meng: Did this reduction come at the cost of making the model less accurate?
Tom: That was the big question, and they used a very strict statistical test called TOST to check it.
Jane: They found that in thirteen out of fifteen different tests, the accuracy was either equivalent or they couldn't prove it had dropped.
Lu: Only two cases actually showed a non-equivalent degradation, and those were specifically on the DeepSeek-R1-Distill-Llama-70B model.
Meng: So for most models, like the Qwen family, the accuracy stayed totally stable?
Tom: Yes, the Qwen models actually showed some tiny improvements in certain benchmarks.
Jane: And we cannot forget the wall-clock latency, which dropped by about ten point five percent on average.
Lu: That is the real win for anyone running these models in production.
Meng: If I can get a ten percent speedup without losing accuracy, my engineering team is going to love this.
Tom: They also showed that the compressed traces are still completely readable by humans.
Jane: You can still see the "backtracking" and "strategy shifts" through those supertoken signposts.
Lu: It turns the reasoning into a map of logical moves rather than just a wall of text.
Meng: It makes the whole process much more transparent.
Tom: It really feels like they have found a way to optimize the very language of thought.
Conclusion: Tom: We have reached the end of our discussion on "Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens."
Jane: It is such a clever way to make eye reasoning more efficient without losing the human-readable logic.
Lu: I think this opens up a whole new way of thinking about how we design the vocabulary for specialized eye agents.
Meng: From my side, the practical implications for reducing inference costs are massive.
Lalam: This advancement allows the intelligence we build to be more streamlined, which eventually helps human culture process complex information more gracefully.
Tom: Thanks for joining us to break down this incredible paper.
Jane: We will see you all next time for the next big discovery on arXiv.
Lila Sciences · Writer, Inc.
cs.CL
Submitted: 2026-04-29
Updated: 2026-09-14
Comments: Accepted to COLM 2026. Code available at https://github.com/Writer/shorthand-for-thought
Code: https://github.com/Writer/shorthand-for-thought
License: http://creativecommons.org/licenses/by-sa/4.0/
Importance score: 85/100
The gist: The paper "Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens" proposes a vocabulary-level approach to compressing chain-of-thought (CoT) reasoning traces in Large
Key concepts
- Supertokens
- Supertokens are single units created by bundling multiple small, predictable pieces of text into one. Instead of a model generating several individual tokens for a common phrase, it uses one supertoken, which increases efficiency during the reasoning process without losing the underlying meaning.
- Structural vs. Organic Tokens
- Reasoning tokens are split into two groups: structural tokens act as predictable scaffolding, like the phrase "Let us check," while organic tokens represent the core logic, such as variable names, numbers, or mathematical steps. Compression focuses on the low-entropy structural tokens.
- Entropy-Guided Compression
- This approach uses mathematics to identify highly predictable phrases that carry little new information. By targeting these low-entropy patterns through BPE merges on the model's own reasoning traces, the system can bundle them into more efficient tokens to save processing time.
Terminology
Summary
The paper Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens
proposes a vocabulary-level approach to compressing chain-of-thought (CoT) reasoning traces in Large Language Models. The authors observe that "reasoning tokens split into two functional types: low-entropy structural tokens (recurring phrases that scaffold the reasoning process) and higher-entropy organic tokens (problem-specific content that drives toward a solution). Structural tokens are described as
recurring, formulaic phrases that scaffold reasoning, such as backtracking cues ('Wait, hold on'), verification phrases ('let us verify'), and strategy shifts ('Let me try'), which organize the reasoning flow but carry little problem-specific information, while organic tokens are
problem-specific content that drives toward a solution, including mathematical expressions, variable bindings, intermediate results, and reasoning direction changes."
The method is a "simple, model-agnostic compression pipeline: apply cross-word BPE merges on a model's own reasoning traces to derive supertokens that capture frequent structural patterns, then teach the model to adopt them via supervised fine-tuning. Specifically, the pipeline involves:
(1) collect reasoning traces, (2) derive supertokens via BPE merges, (3) fine-tune to adopt them. The authors apply SuperBPE to a corpus of reasoning traces, deriving
a small set of reasoning supertokens that are appended to the base vocabulary of a pretrained LLM. They use a structural filter restricting merges to four surface patterns:
capitalized phrase-initial spans (e.g. 'The answer is'), punctuation-plus-newline continuations, comma-led continuations, and space-prefixed single digits (multi-digit concatenations excluded). After filtering, they retain the top-K=250 merges ranked by corpus frequency. New token embeddings are initialized
by averaging the embeddings of their constituent tokens."
The information-theoretic motivation is formalized using conditional entropy: "Given a reasoning trace T = (t1,..., tn) generated by a model with parameters θ, the conditional entropy at position i is: Hθ(ti t<i) = −∑v∈V Pθ(ti = v t<i) log Pθ(ti = v t<i). The theoretical compression ceiling is given by ∆ = ρ·(1 − hM/log2V), where ρ is the fraction of tokens participating in merges and hM is the mean entropy within merge spans. The entropy validation shows that
continuation-token entropy decreases with merge length: length-2 merges already show a 33% entropy reduction relative to non-merge tokens, and merges of length ≥16 reach 79% reduction, confirming that
the merges target sequential redundancy."
The experimental setup evaluates three model families and five mathematical reasoning benchmarks
(AIME'24, AIME'25, MATH-500, Minerva, OlympiadBench) in a zero-shot setting. The training uses a per-device batch size of 1 with 16 gradient accumulation steps (effective batch size 128), a learning rate of 2 × 10−4 with cosine decay and 3% linear warmup, for 1 epoch.
The layer-unfreezing configuration uses N=3 first, M=1 last transformer layers, with backbone learning rate 2 × 10−4,
held fixed across all 15 model–benchmark cells.
The main results show that SFT on a retokenized dataset consistently drives supertoken adoption across all three model families and five benchmarks, reducing reasoning trace length by 6–12% on average,
with an overall average compression of 8.1%. The authors note that "the reduction is attributable to the expanded vocabulary rather than to any behavioral shortening induced by fine-tuning: the SFT (no supertoken) control, trained on the same data with identical hyperparameters but no vocabulary expansion, leaves token length essentially unchanged (−0.47% on average, TOST-equivalent to zero). Wall-clock latency reductions
generally track or exceed token count reductions," with an average latency drop of 10.5% across the three models (QwQ-32B −14.1%, Qwen3-30B-A3B −13.1%, DeepSeek-R1-Distill-Llama-70B −4.4%).
For accuracy, the authors use a two one-sided test (TOST) at a single equivalence margin of ±2 pp
and report that "of the 15 model-benchmark cells, 2/15 demonstrate accuracy equivalence at ±2 pp (QwQ-32B and Qwen3-30B-A3B on MATH-500), 11/15 are inconclusive (90% CI straddles a margin bound, predominantly AIME at N=30), and 2/15 fail equivalence with a real accuracy loss: DeepSeek-R1-Distill-Llama-70B on MATH-500 (−3.1 pp, 90% CI [−5.18, −1.02]) and on OlympiadBench (−2.9 pp, 90% CI [−4.97, −0.83])."
Beyond compression, the authors find that learned supertokens often align with interpretable reasoning moves such as backtracking, verification, and strategy shifts.
They classify supertokens into a nine-category taxonomy (Backtracking, Hedging, Verification, Counterargument, Strategy Shift, Problem Reference, Sequencing, Reasoning, Computation) using deterministic, rule-based classifier that operates on the decoded text of each supertoken merge.
The structural analysis reveals that correct traces show more recovery and verification patterns, while incorrect traces show more repeated hedging and unresolved counterarguments.
Specifically, incorrect traces show elevated rates of confusion cycles: problem-reference→hedging (2.1× over-represented), counterargument→problem-reference (2.0×), and hedging→hedging (1.4×),
while correct traces show elevated rates of productive recovery: problem-reference→strategy-shift (3.0×), verification→strategy-shift (2.1×), and reasoning→verification (1.6×).
The authors acknowledge limitations: Our approach achieves more modest compression (8.1%) compared to content-level CoT compression methods that report 30–70% reductions,
but argue this is a deliberate design choice since by operating at the vocabulary level, we compress the tokenization of reasoning rather than its content, preserving the complete human-readable trace.
They also note the evaluation is restricted to mathematical reasoning
and that the learned vocabulary may not transfer cleanly across tokenizer families.
The accuracy degradation on DeepSeek-R1-Distill-Llama-70B is discussed, with the authors noting that "cleanly disentangling tokenizer-family effects from model-specific effects requires evaluating a second Llama-family model (e.g., R1-Distill-Llama-8B or Llama-3.3-70B-Instruct), which is the priority follow-up we plan to run."
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in an AI system:
What I can build: A post-training module that:
-
Collects the model's own reasoning traces (e.g., from OpenThoughts3 or similar)
-
Applies cross-word BPE merges with the paper's four structural filters (capitalized phrase-initial spans, punctuation-plus-newline, comma-led continuations, space-prefixed single digits)
-
Appends the top-250 merged supertokens to the vocabulary
-
Initializes new embeddings as the mean of constituent token embeddings
-
Fine-tunes only the embedding layer, LM head, and first-3/last-1 transformer layers (or LoRA with r=16, α=32) on retokenized data
Resulting capability: The model produces reasoning traces that are 8.1% shorter on average (6–12% across benchmarks), with wall-clock latency reduced by 10.5%, while maintaining accuracy within ±2 pp on 13/15 model–benchmark cells. This directly reduces inference cost per query without removing reasoning content.
Abstract
Reasoning in Large Language Models incurs significant inference-time compute, yet the token-level information structure of reasoning traces remains underexplored. We observe that reasoning tokens split into two functional types: low-entropy structural tokens (recurring phrases that scaffold the reasoning process) and higher-entropy organic tokens (problem-specific content that drives toward a solution). This asymmetry motivates a simple, model-agnostic compression pipeline: apply cross-word BPE merges on a model's own reasoning traces to derive supertokens that capture frequent structural patterns, then teach the model to adopt them via supervised fine-tuning. Across three model families and five mathematical reasoning benchmarks, our approach shortens reasoning traces by 8.1% on average; under a TOST equivalence analysis at a +/- 2pp margin, accuracy is equivalent or inconclusive on 13/15 model -- benchmark cells (2 pass equivalence, 11 inconclusive, predominantly AIME at N=30, with non-equivalent degradation on 2/15 cells (DeepSeek-R1-Distill-Llama-70B on MATH-500 and OlympiadBench). Beyond compression, learned supertokens often align with interpretable reasoning moves such as backtracking, verification, and strategy shifts. This enables a compact structural analysis of reasoning traces: correct traces show more recovery and verification patterns, while incorrect traces show more repeated hedging and unresolved counterarguments. We release the full pipeline as open-source code.
Sources
- Thought Anchors: Which LLM Reasoning Steps Matter?
- Evaluating Large Language Models Trained on Code
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
- CtrlCoT: Dual-Granularity Chain-of-Thought Compression for Controllable Reasoning
- ConMax: Confidence-Maximizing Compression for Efficient Chain-of-Thought Reasoning
- Modeling Hierarchical Thinking in Large Reasoning Models
- R1-Compress: Long Chain-of-Thought Compression via Chunk Compression and Search
- Qwen3 Technical Report
- Not All Tokens Are What You Need In Thinking
- Instruction-Following Evaluation for Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering