EntropyMoE: Entropy-Aware Sparse Expert Routing for Tokenizer-Free LLMs
Bo Liu, Muxuan Yu, Yu Zhang, Pengfei Gao, Yongping Zhang
University of Bristol · School of Automation Science and Electrical Engineering, Beihang University · University of Manchester
cs.AI
Submitted: 2026-07-31
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 46/100
The gist: EntropyMoE is a "patch-native sparse expert architecture" designed for "tokenizer-free language modeling" that addresses the limitation in existing byte-patch architectures where "existing bytepatch
Terminology
Summary
EntropyMoE is a patch-native sparse expert architecture
designed for tokenizer-free language modeling
that addresses the limitation in existing byte-patch architectures where existing bytepatch architectures still apply the same dense feed-forward computation to every patch.
To resolve this, EntropyMoE replaces the dense feed-forward modules in the global patch Transformer with Top-K expert layers,
specifically utilizing Top-2 expert layers
where each dynamic patch serves as the basic unit of expert routing.
The architecture is built on a BLT-style byte-patch pipeline
where predictable regions are represented by longer patches, whereas uncertain regions are divided into shorter ones.
The core innovation is that the router selects experts directly from patch entropy, using the same granularity signal that underlies dynamic patch construction to organize sparse computation.
Unlike conventional Mixture-of-Experts (MoE) models that route based on high-dimensional hidden states, its router maps only the scalar entropy associated with a patch to expert logits through a learned affine transformation,
expressed as s i = a H i + b. Consequently, the patch hidden state is deliberately excluded from expert selection but remains the input to the selected feed-forward experts,
allowing expert assignment [to be] organized by local uncertainty, while semantic and contextual information is preserved within expert computation.
To account for the varying sizes of dynamic patches, EntropyMoE utilizes byte-weighted routing mass
for workload accounting, defined as rho e = sum i L i p i,e. This mechanism ensures that expert exposure [is measured] using the byte-weighted routing mass
rather than patch counts, which does not measure how much byte-level data is associated with each expert.
Experimental evaluations using the BLT-1B backbone show that EntropyMoE achieves the lowest held-out bits-per-byte (BPB) among matched dense and sparse baselines while maintaining comparable downstream accuracy.
Specifically, it outperformed Dense BLT, Hidden-only MoE, and Hidden+Entropy MoE
in modeling quality. In downstream tasks, it demonstrated a statistically supported improvement on HellaSwag
and maintained competitive performance across the downstream suite.
Routing analysis indicates that EntropyMoE learns a markedly more selective and entropy-responsive routing structure
compared to hidden-state-based routers. The model directs most assignments to a compact expert subset,
where the two most active experts account for 89.1%
of the assignment mass. Furthermore, it exhibits substantially greater divergence between low- and high-entropy routing distributions than the other sparse variants,
proving that patch entropy provides a strong signal for differentiating expert utilization across patch regimes.
In terms of efficiency, the scalar router reduces routing parameters by three orders of magnitude
(400 effective parameters compared to more than 4 times 10 5 for hidden-state controls) and provides the best iteration time and valid-byte throughput among the sparse models.
However, the researchers note that dispatch, communication, and memory movement keep sparse execution slower than Dense BLT.
Improvements for AI systems
1. Entropy-Driven Scalar Routing for Ultra-Lightweight MoE
-
The Improvement: Replace high-dimensional, hidden-state-based routing mechanisms with a scalar-based affine transformation (s i = a H i + b) that utilizes local patch entropy as the sole routing signal.
-
What the improved system can do: It can execute Mixture-of-Experts (MoE) routing with a 1,000x reduction in routing parameters (e.g., 400 vs. 4 times 10 5), drastically lowering the memory overhead of the router. This allows the model to perform highly selective, entropy-responsive expert assignment—directing complex, high-uncertainty data to specialized experts—without the computational cost of processing high-dimensional vectors for every routing decision.
2. Unified Dynamic Patching and Sparse Expert Assignment
-
The Improvement: Integrate a BLT-style byte-patch pipeline with a Top-K sparse expert architecture, where the granularity of the patch (determined by uncertainty) and the expert routing (determined by entropy) are driven by the same signal.
-
What the improved system can do: It can function as a
tokenizer-free
language model that processes raw bytes. The system will automatically adapt its computational resolution in real-time: it will use long patches for predictable, low-entropy text (reducing computation) and short patches for complex, high-entropy text (increasing resolution), while simultaneously routing those specific patches to the most relevant experts.
3. Byte-Weighted Workload Balancing
-
The Improvement: Implement
byte-weighted routing mass
(rho e = sum i L i p i,e) for expert load balancing instead of traditional patch-count-based balancing. -
What the improved system can do: It can ensure equitable expert training and utilization by accounting for the actual volume of byte-level data processed. This prevents experts from being biased toward either
easy
long patches orhard
short patches, ensuring that expert exposure is measured by the actual information density processed, leading to lower bits-per-byte (BPB) and higher modeling quality.
4. Hardware-Aware Sparse Dispatching Kernels
-
The Improvement: Develop specialized communication and memory-movement kernels designed specifically to handle the irregular data patterns created by entropy-based dynamic patching.
-
What the improved system can do: It can bridge the efficiency gap between sparse and dense architectures. By optimizing the dispatch and communication overhead identified in the paper, the system can provide the superior modeling accuracy and low-parameter footprint of EntropyMoE while achieving the high valid-byte throughput and iteration speeds of dense models.
Abstract
Recent byte-level large language models (LLMs) have made tokenizer-free modeling increasingly competitive by grouping bytes into dynamically sized patches. However, existing byte-patch architectures still apply the same dense feed-forward computation to every patch. This uniform computation cannot adapt model capacity to variations in patch semantics and granularity. We address this limitation with EntropyMoE, a Mixture-of-Experts (MoE) architecture designed for dynamic byte patches. EntropyMoE replaces the dense feed-forward modules in the global patch Transformer with Top-K expert layers. Each dynamic patch serves as the basic unit of expert routing, and its byte coverage determines its contribution to workload accounting. The router selects experts directly from patch entropy, using the same granularity signal that underlies dynamic patch construction to organize sparse computation. Patch entropy and length jointly define the feature space for regulating expert specialization. Experiments show that EntropyMoE achieves the lowest held-out bits-per-byte among matched dense and sparse baselines while maintaining comparable downstream accuracy. These results establish patch entropy as an effective routing coordinate for sparse conditional computation and extend Mixture-of-Experts modeling beyond tokenizer-based representations.
Sources
- $\phi$-Balancing for Mixture-of-Experts Training
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
- DeepSeek-V3 Technical Report
- The Llama 3 Herd of Models
- Gemma 3 Technical Report
- Training Compute-Optimal Large Language Models
- Mixtral of Experts
- Fast Byte Latent Transformer
- FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models
- Scaling Laws for Neural Language Models
- Three Phases of Expert Routing: How Load Balance Evolves During Mixture-of-Experts Training
- 2 OLMo 2 Furious
- Byte Latent Transformer: Patches Scale Better Than Tokens
- Qwen2.5 Technical Report
- SpaceByte: Towards Deleting Tokenization from Large Language Modeling
- LLaMA: Open and Efficient Foundation Language Models
- MambaByte: Token-free Selective State Space Model
- Qwen3 Technical Report
- MEGABYTE: Predicting Million-byte Sequences with Multiscale Transformers
- SoftMoE: Soft Differentiable Routing for Mixture-of-Experts in LLMs
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection