You Can Learn Tokenization End-to-End with Reinforcement Learning

arXiv:2602.13940 · cs.LG, cs.AI · Submitted 2026-02-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "You Can Learn Tokenization End-to-End with Reinforcement Learning".

Jane: The paper was written by Sam Dauncey, Roger Wattenhofer and Department of Electrical Engineering, ETH Zürich from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary: Tom: So, we established that "You Can Learn Tokenization End-to-End with Reinforcement Learning" is about giving the AI a much smarter, adaptive way to chop up text. Now, let's talk about what the paper actually summarizes regarding this process.

Jane: If I understand correctly, the core idea is that instead of having separate tokenizers feeding into an LLM, the RL agent learns to generate tokens that are optimal for maximizing performance on a given task.

Tom: It’s not just about making smaller pieces or bigger pieces; it's about making the *right* pieces for the context. Jane, can you elaborate on how that reward mechanism works in practice?

Jane: Well, when we talk about the summary, they show that the agent is rewarded based on how well the resulting sequence of tokens allows a downstream task—say, answering questions—to be performed successfully.

Meng: That makes perfect sense; the tokenization isn't an academic exercise; it's directly tied to performance metrics, which is what any good engineer needs to see.

Lu: It suggests that the model learns to predict not just tokens, but *useful* tokens—tokens that carry maximum semantic weight for the overall objective, rather than just being statistically common.

Lalam: From a broader cultural view, this means AI won't struggle with ambiguities simply because of how we chunk text; it will understand the intended meaning even if the optimal segmentation is unconventional.

Tom: But I’m picturing a system that is constantly self-correcting its own input format, which sounds incredibly powerful and complex to train.

Jane: It is complex, but the paper frames it as an iterative process where failure in one task helps refine the tokenization strategy for all future tasks.

Meng: So, if we could implement this, we wouldn't need a fixed vocabulary size or a static set of rules that might break down when encountering highly specialized jargon.

Lu: That adaptability is the breakthrough; it means the system can handle domain shifts—say, moving from medical transcripts to legal documents—without needing manual recalibration of the tokenization layer.

Lalam: Imagine how much more robust human-computer interaction would be if the underlying language processing could adapt its vocabulary and structure based on the context of conversation itself.

Improvements/Methods: Tom: We’ve talked about what this approach is, and we've seen that it uses RL to improve tokenization. Now, let's focus on what improvements the paper suggests over existing methods, because that's where the real engineering novelty lies.

Jane: I

Paper discussion segment 3: Tom: So, we’ve heard how this paper introduces a powerful way to teach AI models how to chunk text using reinforcement learning. The core idea is that instead of relying on fixed rules like BPE, the model learns to find the most semantically useful boundaries based on performance.

Jane: And that's a huge step forward because, as we know, traditional tokenization is a static process; it’s just a hardcoded compression step before the LLM even starts working. This method allows for dynamic learning of those boundaries.

Meng: That dynamic aspect is what I’m most interested in from an engineering viewpoint—the ability to handle unforeseen text variations without needing manual rule updates, right? How does the system actually know that a boundary should be placed at a whitespace or a logical break, and not just anywhere?

Lu: It's because they use something called a score function estimator which gives the model better theoretical guarantees than the straight-through methods we’ve seen before. The AI is directly optimizing for minimizing loss based on the expected loss.

Lalam: That capability to adapt is so important because language itself evolves, and fixed tokenization struggles with that; it prevents us from understanding nuanced or complex language structures as they emerge in society.

Tom: Exactly, Lu, it’s not just about statistical frequency anymore; it's about semantic weight. But how do we make this theoretically superior score function practical for training?

Jane: That's where the reinforcement learning comes in, specifically by using techniques like time discounting and advantage estimation to reduce the high variance that comes with those scores.

Meng: So, they’re not just guessing at boundaries; they’re being smart about when to place them based on a calculated benefit relative to the expected loss?

Lu: Precisely, Meng; it's a sophisticated form feedback loop that helps guide the policy pi theta toward optimal tokenization.

Lalam: And this ability, if we can scale it well, allows us to create models that are much more robust across different domains and vastly improve how diverse communities interact with AI.

Tom: It sounds like they’ve found a way to make the theoretical ideal of end-to-end learning practically feasible.

Jane: It really is a blend of theory and practical RL optimization.

Conclusion: Tom: So, to wrap up our deep dive into "You Can Learn Tokenization End-to-End with Reinforcement Learning," it really boils down to this: we're looking at a massive shift away from hand-engineered NLP components.

Jane: Exactly. What's beautiful about this work is that it shows we don't need to teach the model how to break up words using pre-defined rules; the model can learn that process itself, just by training on the whole sequence.

Lu: But think about what that means for general AI! If tokenization becomes an emergent property of the language model, it suggests a fundamental unification of representation—that we're moving toward a single, unified understanding of information flow rather than stacking separate modules.

Meng: I hear the excitement in your voice, Lu, but practically speaking, if we remove that intermediate step of explicit tokenization rules or pre-processing layers entirely, how much compute overhead are we really eliminating in a large-scale production environment?

Lalam: The removal of those rigid boundaries will have a profound effect on human culture. By making the foundational process of language interpretation fluid and learned, AI can assist humanity not just in creating better software, but in cultivating new forms of spontaneous, cross-cultural communication that reflect true cognitive flexibility.

Tom: Lalam’s point about cultural fluidity is huge; it really puts the scope beyond just optimizing performance metrics. We should keep keeping an eye on how this advances the entire field of foundation models.

Jane: It's a truly exciting time for NLP, and I think that learning how to tackle these massive, fundamental design choices is what makes AI so compelling right now.

Lu: I just hope the next wave of research can apply this principle—this end-to-end learning—to modalities beyond text, like video or sensory data.

Meng: If we can solve tokenization this way for language, then applying that structural simplicity to complex multimodal inputs feels like the natural engineering next step.

Lalam: The ability to learn the core structure of data representations, whatever that data is, is ultimately what will help us improve how we connect and communicate with one another.

Tom: And so, we're going to leave you with all that exciting thought about "You Can Learn Tokenization End-to-End with Reinforcement Learning."

Jane: We’ll catch up next time when we tackle another fascinating paper and explore what it means for our collective future.

Curran Associates · Advances in Neural Information Processing Systems

cs.LG, cs.AI

Submitted: 2026-02-15

Updated: 2026-09-14

Code: https://github.com/SamD770/bitter-lesson-tokenization

Project page: https://main-horse.github.io/hnet

Importance score: 73/100

The gist: This paper presents a method for learning tokenization strategies end-to-end using reinforcement learning, aiming to replace the "hardcoded compression step" currently used in Large Language Model

Key concepts

Reinforcement Learning (RL)
The AI agent learns through an iterative process where it is rewarded for success. The reward system is based on how well the resulting sequence of tokens allows a downstream task, such as answering questions, to be performed successfully.
Tokenization
This refers to the process of chopping or segmenting text into smaller pieces (tokens). The paper's method allows this process to be learned dynamically by finding the most semantically useful boundaries, rather than using fixed, pre-defined rules.
End-to-End Learning
This concept involves training an AI model where the tokenization process is not a separate step but is learned as part of the entire sequence. This suggests a unified understanding of information flow, eliminating the need for manual pre-processing layers.

Terminology

Summary

This paper presents a method for learning tokenization strategies end-to-end using reinforcement learning, aiming to replace the hardcoded compression step currently used in Large Language Model (LLM) training pipelines. By treating token boundary placement as a discrete decision optimized via score function estimators, the authors seek to move beyond largely artisanal tokenizer design toward architectures that can autonomously discover semantic boundaries.

The Proposed Architecture

The researchers propose an autoregressive U-net architecture designed to bring the tokenization process inside the architecture and training of an LLM. To satisfy their desideratum for end-to-end architecture, the model moves from the byte level to the token level and back through a specific sequence:

  1. Autoregressively encode input bytes into byte-level representations X.

  2. (Potentially stochastically) predict token boundaries a i from these representations.

  3. Downsample the byte-level representations into token-level representations X'.

  4. Enrich token-level representations with an autoregressive feedforward network.

  5. Upsample into updated byte-level information Y.

  6. Decode the resulting byte-level representations to predict the next bytes y i.

Learning via Score Function Estimation

Because tokenizing a byte stream is an inherently discrete decision, standard gradient-based optimization cannot be applied directly. The authors instead utilize score function estimators, which directly approximate the gradient of the expected loss with respect to the token boundaries. This method offers stronger theoretical guarantees than prior straight-through estimators (STEs) because it optimizes the problem of drawing discrete boundaries directly. To make this approach practicable and reduce high variance, they employ several reinforcement learning techniques:

  • Time discounting: Introducing a discount factor (gamma = 0.99) to decouple advantages from far-away sequence parts.

  • Batch-relative advantages: Centering advantage estimates using multiple values within a single batch.

  • Early exit relative rewards: Subtracting tokenization-independent noise by comparing the model against an early byte-level exit model.

Targeting Downsampling Rates

To prevent the model from opting for a computationally expensive strategy of separating every byte with a token boundary, the authors implement a mechanism to maintain a specific target downsample rate. They apply even pressure across all the logits in the batch using a specialized loss term. To ensure stability and efficient exploration, they utilize:

  • Logit scaling: Using a scaling factor D to make logits approximately uniform at initialization.

  • Softcapping: Applying softcapping to prevent exploding logits during the training process.

Experimental Performance

The method was evaluated at the 100 million parameter scale, where it outperforms prior proposed straight-through estimates, both qualitatively and quantitatively. Despite having no explicit prior structure or inductive bias, the model reliably learns to place boundaries at whitespace-like characters such as newlines and spaces. In benchmark testing on FineWeb and other tasks like PIQA and LAMBADA, the learned policy achieved lower bits-per-byte than previous dynamic methods. Most significantly, the method recovers the performance of BPE-guided downsampling without access to these external priors, demonstrating that it can learn effective compression strategies entirely through optimization.

Improvements for AI systems

Improvements

  1. Integration of a Learned Tokenization Layer via Score Function Estimators: Replace fixed Byte-Pair Encoding (BPE) or hand-crafted subword tokenizers with a stochastic, end-to-end trainable layer within the model architecture. This layer will use score function (policy gradient) estimators rather than straight-through estimators to optimize discrete token boundary decisions directly.

  2. Deployment of an Autoregressive U-Net Architecture with Information Reuse: Implement a hierarchical architecture that encodes raw UTF-8 bytes into latent tokens via downsampling and subsequently upsamples them back to updated byte-level representations. This ensures that high-resolution byte-level information is preserved and reused at the token level.

  3. Implementation of RL Variance Reduction Mechanisms: Incorporate time discounting (gamma = 0.99) to decouple far-away reward signals and utilize an early-exit (pearly) byte-level model to serve as a baseline for rewards. This addresses the reward attribution problem by isolating the impact of specific boundary decisions on subsequent token loss.

  4. Dynamic Downsampling Rate Targeting: Integrate a logit-based pressure mechanism that applies softcapping and target downsampling rates (pi target) to the boundary logits. This prevents the model from defaulting to computationally expensive byte-by-byte tokenization while maintaining optimal compression ratios.

Improved AI System Capabilities

  1. Autonomous Semantic Boundary Discovery: The system will automatically align token boundaries with semantic structures (e.g., whitespace, punctuation, or code syntax like module names) without requiring explicit linguistic priors or manual rule-setting.

  2. Enhanced Cross-Lingual Robustness: The system will provide superior performance for underrepresented and non-English languages by eliminating the bias and glitch token risks inherent in language-specific, artisanal BPE vocabularies.

  3. Superior Compression Efficiency: The model will achieve lower bits-per-byte (BBP) metrics and lower validation loss across natural language and programming languages (e.g., Python) compared to models using fixed tokenization or straight-through estimation.

  4. Optimized Computational Scaling: The system can dynamically adjust its effective downsampling rate to balance the trade-off between sequence length (computational cost) and vocabulary granularity, optimizing FLOPs usage during both training and inference.

Sources

Related papers