Easy to anticipate, hard to compute: boundary dependence finds the computed outputs that entropy patching misses
cs.CL
Submitted: 2026-10-08
Updated: 2026-10-08
Code: https://github.com/nicoveraz/segresearch
Terminology
Sources
- Training Verifiers to Solve Math Word Problems
- ByteSpan: Information-Driven Subword Tokenisation
- Measuring Mathematical Problem Solving With the MATH Dataset
- Dynamic Chunking for End-to-End Hierarchical Sequence Modeling
- MrT5: Dynamic Token Merging for Efficient Byte-level Language Models
- Fast Byte Latent Transformer
- Rho-1: Not All Tokens Are What You Need
- EntropyMoE: Entropy-Aware Sparse Expert Routing for Tokenizer-Free LLMs
- TokEval: A Tokenizer Evaluation Suite
- Efficient Transformers with Dynamic Token Pooling
- Byte Latent Transformer: Patches Scale Better Than Tokens
- OpenWebMath: An Open Dataset of High-Quality Mathematical Web Text
- Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
- Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs
- SpaceByte: Towards Deleting Tokenization from Large Language Modeling
- From Bytes to Ideas: Language Modeling with Autoregressive U-Nets
- Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering