Distilling Token-Trained Models into Byte-Level Models

summary

Video file (mp4)

The gist

The gist: This paper proposes an efficient distillation recipe that converts existing token-trained LLMs into Byte Language Models (BLMs) using a two-stage curriculum, achieving comparable

In short

The paper proposes a two-stage distillation process to convert existing token-trained Large Language Models into Byte Language Models (BLMs). This method aligns byte-level representations with teacher embeddings and then fine-tunes the model entirely in the byte space. The result is a highly efficient BLM that achieves comparable performance using only about 125 billion bytes of training data, significantly lowering the cost of building scalable BLMs.

Key concepts

Progressive Knowledge Distillation
This first stage aligns a student model's representations with a teacher's. It uses three specific losses—embedding alignment, joint distillation, and boundary learning—in sequence to ensure the student understands both the meaning of bytes and where tokens begin and end in the teacher model.
Byte-Level Supervised Fine-Tuning (SFT)
The second stage shifts the model to generate output using only byte representations. The student replaces its token head with byte-specific modules like Dechunk and Decoder, training it directly on next byte prediction loss to enable end-to-end generation within the byte space.
ONE-BYTE LOOKAHEAD ROUTING
This mechanism is part of the H-Net architecture used for byte modeling. It calculates similarity between the current byte's hidden state and the next one to determine a boundary indicator, which helps guide the model in predicting where one byte sequence ends and another begins.
Space Bias
This refers to a problem where the boundary predictor in Stage 1 becomes overly reliant on surface-level patterns rather than deeper semantic context. The authors propose solutions like Trim Data and Whitespace Penalty to mitigate this overfitting.

Terminology used across episodes

This episode discusses

The paper

Distilling Token-Trained Models into Byte-Level Models · Read on arXiv

Zishuo Bao, Jiaqi Leng, Junxiong Wang, Bowen Peng, Yucheng Lu

NYU Shanghai

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Distilling Token-Trained Models into Byte-Level Models".

Tom: The gist: This paper proposes an efficient distillation recipe that converts existing token-trained LLMs into Byte Language Models (BLMs) using a two-stage curriculum,

Jane: First, who's behind it and why it matters.

Paper summary: Tom: We've talked about the core idea of this paper, "Distilling Token-Trained Models into Byte-Level Models". The thesis is that existing token-trained models can be efficiently converted into byte language models without needing massive amounts of data from scratch.

Jane: They achieve this by proposing a two-stage curriculum. First, progressive knowledge distillation which aligns the byte representations with the teacher model's embeddings, and second, byte-level supervised fine-tuning to enable end-to-end generation entirely in the byte space.

Lu: The paper substantiates this recipe over models like Llama, Qwen, and OLMo showing that this approach results in very efficient data usage. They demonstrate that with about one hundred twenty-five B training bytes, the distilled model keeps most of its teacher capabilities <ref:2602.01007#pg1>.

Meng: That efficiency is what matters for practical application. So if you have a large model already trained, this recipe lets you create a smaller byte-level version without needing to re-train from scratch on trillions of bytes.

Lalam: It really focuses on reducing the data requirement significantly, which is crucial since training models from scratch for BLMs typically requires trillions of bytes.

Tom: And the results show that even with this efficiency, they maintain strong performance retention across various benchmarks, like MMLU scores around fifty-one point eight and sixty-eight point five for Llama-three point two 3B and Qwen-three 4B respectively <ref:2602.01007#pg2,51.8 and 68.5 for Llama-3.2 3B and Qwen>.

Jane: So what this means is that building these byte models becomes much more feasible because you can leverage the knowledge already in those larger token-trained models instead of starting from zero.

Lu: The paper’s main contribution is proposing this novel two-stage distillation framework that effectively transforms token-trained LLMs into BLMs, which drastically reduces the data requirement compared to training BLMs from scratch.

Meng: It's a practical recipe for resource conservation in building AI systems that need to operate directly in the byte space.

Lalam: This work substantially lowers the cost of building scalable BLMs by showing a practical way to convert token-trained LLMs into byte-level models while retaining most of their abilities using only one hundred twenty-five B training bytes <ref:2602.01007#pg1>.

Conclusion: Tom: So wrapping up this discussion on "Distilling Token-Trained Models into Byte-Level Models". The authors are Jiaqi Leng, Junxiong Wang, Bowen Peng, and Yucheng Lu.

Jane: They introduced a practical two-stage distillation framework that converts token-trained LLMs into byte language models while preserving most of their capabilities using just about one hundred twenty-five billion bytes of training data.

Lu: The authors validate this approach across multiple model families, including Llama, Qwen, and OLMo.

Meng: It shows that leveraging existing knowledge is a viable strategy for creating smaller byte-level models without needing to train from scratch on massive datasets.

Lalam: This method substantially lowers the cost of building scalable BLMs by showing a practical way to convert token-trained LLMs into byte-level models while retaining most of their abilities using only one hundred twenty-five B training bytes <ref:2602.01007#pg1>.

More episodes

← Home