Distilling Token-Trained Models into Byte-Level Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Distilling Token-Trained Models into Byte-Level Models".
Tom: The gist: This paper proposes an efficient distillation recipe that converts existing token-trained LLMs into Byte Language Models (BLMs) using a two-stage curriculum,
Jane: First, who's behind it and why it matters.
Paper summary: Tom: We've talked about the core idea of this paper, "Distilling Token-Trained Models into Byte-Level Models". The thesis is that existing token-trained models can be efficiently converted into byte language models without needing massive amounts of data from scratch.
Jane: They achieve this by proposing a two-stage curriculum. First, progressive knowledge distillation which aligns the byte representations with the teacher model's embeddings, and second, byte-level supervised fine-tuning to enable end-to-end generation entirely in the byte space.
Lu: The paper substantiates this recipe over models like Llama, Qwen, and OLMo showing that this approach results in very efficient data usage. They demonstrate that with about one hundred twenty-five B training bytes, the distilled model keeps most of its teacher capabilities <ref:2602.01007#pg1>.
Meng: That efficiency is what matters for practical application. So if you have a large model already trained, this recipe lets you create a smaller byte-level version without needing to re-train from scratch on trillions of bytes.
Lalam: It really focuses on reducing the data requirement significantly, which is crucial since training models from scratch for BLMs typically requires trillions of bytes.
Tom: And the results show that even with this efficiency, they maintain strong performance retention across various benchmarks, like MMLU scores around fifty-one point eight and sixty-eight point five for Llama-three point two 3B and Qwen-three 4B respectively <ref:2602.01007#pg2,51.8 and 68.5 for Llama-3.2 3B and Qwen>.
Jane: So what this means is that building these byte models becomes much more feasible because you can leverage the knowledge already in those larger token-trained models instead of starting from zero.
Lu: The paper’s main contribution is proposing this novel two-stage distillation framework that effectively transforms token-trained LLMs into BLMs, which drastically reduces the data requirement compared to training BLMs from scratch.
Meng: It's a practical recipe for resource conservation in building AI systems that need to operate directly in the byte space.
Lalam: This work substantially lowers the cost of building scalable BLMs by showing a practical way to convert token-trained LLMs into byte-level models while retaining most of their abilities using only one hundred twenty-five B training bytes <ref:2602.01007#pg1>.
Conclusion: Tom: So wrapping up this discussion on "Distilling Token-Trained Models into Byte-Level Models". The authors are Jiaqi Leng, Junxiong Wang, Bowen Peng, and Yucheng Lu.
Jane: They introduced a practical two-stage distillation framework that converts token-trained LLMs into byte language models while preserving most of their capabilities using just about one hundred twenty-five billion bytes of training data.
Lu: The authors validate this approach across multiple model families, including Llama, Qwen, and OLMo.
Meng: It shows that leveraging existing knowledge is a viable strategy for creating smaller byte-level models without needing to train from scratch on massive datasets.
Lalam: This method substantially lowers the cost of building scalable BLMs by showing a practical way to convert token-trained LLMs into byte-level models while retaining most of their abilities using only one hundred twenty-five B training bytes <ref:2602.01007#pg1>.
Zishuo Bao, Jiaqi Leng, Junxiong Wang, Bowen Peng, Yucheng Lu
NYU Shanghai
cs.CL
Submitted: 2026-02-01
Updated: 2026-10-05
Code: https://github.com/heavyball-research/DistillBytes
Importance score: 77/100
The gist: The gist: This paper proposes an efficient distillation recipe that converts existing token-trained LLMs into Byte Language Models (BLMs) using a two-stage curriculum, achieving comparable
Key concepts
- Progressive Knowledge Distillation
- This first stage aligns a student model's representations with a teacher's. It uses three specific losses—embedding alignment, joint distillation, and boundary learning—in sequence to ensure the student understands both the meaning of bytes and where tokens begin and end in the teacher model.
- Byte-Level Supervised Fine-Tuning (SFT)
- The second stage shifts the model to generate output using only byte representations. The student replaces its token head with byte-specific modules like Dechunk and Decoder, training it directly on next byte prediction loss to enable end-to-end generation within the byte space.
- ONE-BYTE LOOKAHEAD ROUTING
- This mechanism is part of the H-Net architecture used for byte modeling. It calculates similarity between the current byte's hidden state and the next one to determine a boundary indicator, which helps guide the model in predicting where one byte sequence ends and another begins.
- Space Bias
- This refers to a problem where the boundary predictor in Stage 1 becomes overly reliant on surface-level patterns rather than deeper semantic context. The authors propose solutions like Trim Data and Whitespace Penalty to mitigate this overfitting.
Terminology
Summary
The gist: This paper proposes an efficient distillation recipe that converts existing token-trained LLMs into Byte Language Models (BLMs) using a two-stage curriculum, achieving comparable performance with only approximately 125B bytes of training data.
Proposed Distillation Recipe
The paper introduces a novel distillation recipe that follows a two-stage curriculum to convert existing token-trained LLMs into BLMs while retaining comparable capabilities. This recipe is structured into (1) Progressive Knowledge Distillation, which aligns byte-level representations with the embeddings of the token-trained teacher model, and (2) Byte-Level Supervised Fine-Tuning, which enables end-to-end generation entirely in the byte space. The central challenge addressed is the boundary mismatch between token-based models and byte models.
Stage 1: Progressive Knowledge Distillation
Stage 1 focuses on progressive representation alignment and boundary learning. This stage involves three specific loss objectives to align the student model’s embedding space, semantic probabilities, and boundary decisions with those of the teacher model.
-
Embedding Alignment (Lalign): This objective minimizes the difference between the hidden state of the student model encoder at the boundary byte eˆk and the static embedding of the teacher model ek.
-
Joint Distillation (Ldistill): This objective minimizes the KL divergence between the student model’s and teacher model’s output distributions by utilizing the teacher model’s segmentation boundaries to synchronize sequence lengths.
-
Boundary Learning (Lboundary): This objective trains the Routing Module to predict token boundaries via binary classification on each byte, minimizing the loss defined as Lboundary.
The framework adopts a sequential curriculum rather than a holistic objective to ensure robustness and generalizability, optimizing Encoder Alignment first, then Joint Distillation, and finally Boundary Learning.
Stage 2: Byte-Level Supervised Fine-Tuning (SFT)
Building on Stage 1, Stage 2 transitions the model to generate entirely in the byte space. The student model replaces the token LM head with Dechunk and Decoder modules and is trained using the Next Byte Prediction loss. This process involves two steps: Head Adaptation, training only the newly initialized Dechunk and Decoder modules, followed by end-to-end fine-tuning where all parameters are unfrozen to optimize the entire model.
Architectural Modifications for Byte Modeling
The student model is based on H-Net, a U-Net-like architecture crafted for multigranularity sequence modeling. Key architectural adaptations include the ONE-BYTE LOOKAHEAD ROUTING mechanism, which computes the similarity between the current byte hidden state xcurr and the next byte hidden state xnext to derive a boundary indicator b = I(p ≥ 0.5). For decoding, two strategies are explored: Joint Boundary Prediction (JBP), which encodes boundary information directly into the output vocabulary, and Multi-Byte Prediction (MBP), which introduces an auxiliary language modeling head to predict the next-next byte to simulate lookahead behavior during inference. Furthermore, a Dechunk module employs a dynamic pooling strategy by revealing chunk representations only at the final byte of the k-th chunk to preserve autoregressive properties while providing maximal context at chunk boundaries.
Empirical Validation and Results
The distillation recipe is validated across multiple model families, including Llama, Qwen, and OLMo. Empirical results demonstrate unprecedented data efficiency, with approximately 125 B training bytes in total required for the distilled model. For instance, on the MMLU benchmark, the method achieves scores of 51.8 and 68.5 for Llama-3.2 3B and Qwen-3 4B respectively, retaining over 92% of the original performance. The framework shows superior performance retention compared to baselines like BLT and H-Net Distill, achieving lower average drop across various benchmarks.
Robustness and Further Enhancements
The distilled Stage 1 models demonstrate a significant level of intrinsic resilience against character-level perturbations. However, the analysis reveals that the boundary predictor is heavily overfitted to surface patterns, suggesting a Space Bias. Proposed mitigation strategies include Trim Data, which removes all whitespaces to force reliance on character sequences alone, and Whitespace Penalty, which introduces a penalty term in the boundary loss function to discourage boundaries at whitespace positions. Fine-tuning on perturbed benchmarks shows that the byte-level Llama distill model exhibits a significantly higher adaptation ceiling than its subword counterpart, achieving the highest overall Robustness Score (69.93) after 10 epochs of training on noisy data. The findings also indicate that the Mamba2 encoder offers high fidelity compared to standard Transformer encoders in aligning fine-grained byte sequences with high-level semantic embeddings.
Comparison with Related Work
The framework diverges from prior approaches like Bolmo by decomposing the process into a sequential curriculum, explicitly separating representation alignment and capability transfer from boundary learning. The paper also highlights that the performance of decoupled routing modules suffers a notable decline compared to integrated ones, suggesting that token boundaries are heavily dependent on high-level semantic and syntactic contexts provided by a fully developed encoder. Moreover, the study confirms that employing a significantly larger teacher model does not offer a distinct advantage in this distillation context, as standard setups with similarly scaled models consistently produce better outcomes. The final results show that while the distilled model is competitive, there remains room for improvement to match the performance of models trained with different objectives or larger computational budgets.
Conclusion
In this paper, a practical two-stage distillation framework is proposed that converts pretrained token-based LLMs into byte-level models while preserving most of their capabilities. This method is effective across multiple model families, including Llama, Qwen, and OLMo, and achieves competitive performance using only 125B training bytes. The distillation process successfully enables the H-Net to utilize its byte-to-byte architecture to bypass the fragility of the subword boundaries. This work substantially lowers the cost of building scalable BLMs. The distillation process successfully enables the H-Net to utilize its byte-to-byte architecture to bypass the fragility of the subword boundaries. This work substantially lowers the cost of building scalable BLMs. The distillation process successfully enables the H-Net to utilize its byte-to-byte architecture to bypass the fragility of the subword boundaries. This work substantially lowers the cost of building scalable BLMs. The distillation process successfully enables the H-Net to utilize its byte-to-byte architecture to bypass the fragility of the subword boundaries. This work substantially lowers the cost of building scalable BLMs. The distillation process successfully enables the H-Net to utilize its byte-to-byte architecture to bypass the fragility of the subword boundaries. This work substantially lowers the cost of building scalable BLMs. The distillation process successfully enables the H-Net to utilize its byte-to-byte architecture to bypass the fragility of the subword boundaries. This work substantially lowers the cost of building scalable BLMs. The distillation process successfully enables the H-Net to utilize its byte-to-byte architecture to bypass the fragility of the subword boundaries.
Improvements for AI systems
- Bold header: Byte-Level Model Creation via Two-Stage Distillation
The improved AI system can convert existing token-trained LLMs into byte-level models using a two-stage curriculum,
which drastically reduces data requirements, achieving competitive performance with approximately 125B bytes of training data.
- Bold header: Progressive Knowledge Distillation for Representation Alignment
The system implements Stage 1, Progressive Knowledge Distillation,
utilizing three loss objectives—embedding alignment (Lalign), joint distillation (Ldistill), and boundary learning (Lboundary)—to align byte-level representations with the embeddings of the token-trained teacher model.
- Bold header: Dynamic Boundary Learning for Autonomous Tokenization
The system incorporates a one-byte lookahead mechanism
in the routing module, which is trained via the Boundary Learning Objective (Lboundary)
to predict token boundaries, enabling the student model to dynamically group raw bytes into chunks and generate in the token space.
- Bold header: Robustness Enhancement via Post-Training Optimization
The system can be enhanced using post-training techniques like GOLD integrated with On-Policy Distillation, which leverages the exceptionally high boundary detection accuracy
of the Stage 1 model to further align its segmentation trajectory.
Sources
- PIQA: Reasoning about Physical Commonsense in Natural Language
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- Multiscale Byte Language Models -- A Hierarchical Architecture for Causal Million-Length Sequence Modeling
- Measuring Massive Multitask Language Understanding
- Distilling the Knowledge in a Neural Network
- Dynamic Chunking for End-to-End Hierarchical Sequence Modeling
- Finetuning Pretrained Transformers into RNNs
- Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering
- Universal Cross-Tokenizer Distillation via Approximate Likelihood Matching
- 2 OLMo 2 Furious
- Olmo 3
- WinoGrande: An Adversarial Winograd Schema Challenge at Scale
- Qwen3 Technical Report
- HellaSwag: Can a Machine Really Finish Your Sentence?
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering