EntropyMoE: Entropy-Aware Sparse Expert Routing for Tokenizer-Free LLMs

arXiv:2608.06398 · cs.AI · Submitted 2026-07-31 · Read on arXiv

Bo Liu, Muxuan Yu, Yu Zhang, Pengfei Gao, Yongping Zhang

University of Bristol · School of Automation Science and Electrical Engineering, Beihang University · University of Manchester

cs.AI

Submitted: 2026-07-31

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 46/100

The gist: EntropyMoE is a "patch-native sparse expert architecture" designed for "tokenizer-free language modeling" that addresses the limitation in existing byte-patch architectures where "existing bytepatch

Terminology

Summary

EntropyMoE is a patch-native sparse expert architecture designed for tokenizer-free language modeling that addresses the limitation in existing byte-patch architectures where existing bytepatch architectures still apply the same dense feed-forward computation to every patch. To resolve this, EntropyMoE replaces the dense feed-forward modules in the global patch Transformer with Top-K expert layers, specifically utilizing Top-2 expert layers where each dynamic patch serves as the basic unit of expert routing.

The architecture is built on a BLT-style byte-patch pipeline where predictable regions are represented by longer patches, whereas uncertain regions are divided into shorter ones. The core innovation is that the router selects experts directly from patch entropy, using the same granularity signal that underlies dynamic patch construction to organize sparse computation. Unlike conventional Mixture-of-Experts (MoE) models that route based on high-dimensional hidden states, its router maps only the scalar entropy associated with a patch to expert logits through a learned affine transformation, expressed as s i = a H i + b. Consequently, the patch hidden state is deliberately excluded from expert selection but remains the input to the selected feed-forward experts, allowing expert assignment [to be] organized by local uncertainty, while semantic and contextual information is preserved within expert computation.

To account for the varying sizes of dynamic patches, EntropyMoE utilizes byte-weighted routing mass for workload accounting, defined as rho e = sum i L i p i,e. This mechanism ensures that expert exposure [is measured] using the byte-weighted routing mass rather than patch counts, which does not measure how much byte-level data is associated with each expert.

Experimental evaluations using the BLT-1B backbone show that EntropyMoE achieves the lowest held-out bits-per-byte (BPB) among matched dense and sparse baselines while maintaining comparable downstream accuracy. Specifically, it outperformed Dense BLT, Hidden-only MoE, and Hidden+Entropy MoE in modeling quality. In downstream tasks, it demonstrated a statistically supported improvement on HellaSwag and maintained competitive performance across the downstream suite.

Routing analysis indicates that EntropyMoE learns a markedly more selective and entropy-responsive routing structure compared to hidden-state-based routers. The model directs most assignments to a compact expert subset, where the two most active experts account for 89.1% of the assignment mass. Furthermore, it exhibits substantially greater divergence between low- and high-entropy routing distributions than the other sparse variants, proving that patch entropy provides a strong signal for differentiating expert utilization across patch regimes.

In terms of efficiency, the scalar router reduces routing parameters by three orders of magnitude (400 effective parameters compared to more than 4 times 10 5 for hidden-state controls) and provides the best iteration time and valid-byte throughput among the sparse models. However, the researchers note that dispatch, communication, and memory movement keep sparse execution slower than Dense BLT.

Improvements for AI systems

1. Entropy-Driven Scalar Routing for Ultra-Lightweight MoE

  • The Improvement: Replace high-dimensional, hidden-state-based routing mechanisms with a scalar-based affine transformation (s i = a H i + b) that utilizes local patch entropy as the sole routing signal.

  • What the improved system can do: It can execute Mixture-of-Experts (MoE) routing with a 1,000x reduction in routing parameters (e.g., 400 vs. 4 times 10 5), drastically lowering the memory overhead of the router. This allows the model to perform highly selective, entropy-responsive expert assignment—directing complex, high-uncertainty data to specialized experts—without the computational cost of processing high-dimensional vectors for every routing decision.

2. Unified Dynamic Patching and Sparse Expert Assignment

  • The Improvement: Integrate a BLT-style byte-patch pipeline with a Top-K sparse expert architecture, where the granularity of the patch (determined by uncertainty) and the expert routing (determined by entropy) are driven by the same signal.

  • What the improved system can do: It can function as a tokenizer-free language model that processes raw bytes. The system will automatically adapt its computational resolution in real-time: it will use long patches for predictable, low-entropy text (reducing computation) and short patches for complex, high-entropy text (increasing resolution), while simultaneously routing those specific patches to the most relevant experts.

3. Byte-Weighted Workload Balancing

  • The Improvement: Implement byte-weighted routing mass (rho e = sum i L i p i,e) for expert load balancing instead of traditional patch-count-based balancing.

  • What the improved system can do: It can ensure equitable expert training and utilization by accounting for the actual volume of byte-level data processed. This prevents experts from being biased toward either easy long patches or hard short patches, ensuring that expert exposure is measured by the actual information density processed, leading to lower bits-per-byte (BPB) and higher modeling quality.

4. Hardware-Aware Sparse Dispatching Kernels

  • The Improvement: Develop specialized communication and memory-movement kernels designed specifically to handle the irregular data patterns created by entropy-based dynamic patching.

  • What the improved system can do: It can bridge the efficiency gap between sparse and dense architectures. By optimizing the dispatch and communication overhead identified in the paper, the system can provide the superior modeling accuracy and low-parameter footprint of EntropyMoE while achieving the high valid-byte throughput and iteration speeds of dense models.

Abstract

Recent byte-level large language models (LLMs) have made tokenizer-free modeling increasingly competitive by grouping bytes into dynamically sized patches. However, existing byte-patch architectures still apply the same dense feed-forward computation to every patch. This uniform computation cannot adapt model capacity to variations in patch semantics and granularity. We address this limitation with EntropyMoE, a Mixture-of-Experts (MoE) architecture designed for dynamic byte patches. EntropyMoE replaces the dense feed-forward modules in the global patch Transformer with Top-K expert layers. Each dynamic patch serves as the basic unit of expert routing, and its byte coverage determines its contribution to workload accounting. The router selects experts directly from patch entropy, using the same granularity signal that underlies dynamic patch construction to organize sparse computation. Patch entropy and length jointly define the feature space for regulating expert specialization. Experiments show that EntropyMoE achieves the lowest held-out bits-per-byte among matched dense and sparse baselines while maintaining comparable downstream accuracy. These results establish patch entropy as an effective routing coordinate for sparse conditional computation and extend Mixture-of-Experts modeling beyond tokenizer-based representations.

Sources

Related papers