MultiHashFormer: Hash-based Generative Language Models
cs.CL, cs.AI, cs.LG
Submitted: 2026-06-26
Updated: 2026-08-28
Comments: Accepted at EMNLP 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Language models (LMs) represent tokens using embedding matrices that scale linearly with the vocabulary size.
Terminology
Abstract
Language models (LMs) represent tokens using embedding matrices that scale linearly with the vocabulary size. To constrain the parameter footprint, prior work proposes hashing many tokens into a single vector within encoder-only models. While this offers parameter efficiency, many-to-one collisions prevent its use in causal LMs. In this paper, we propose MultiHashFormer, a new framework that allows hash-based autoregression. Each token is represented as a unique hash signature, a short sequence of discrete hash IDs, generated by multiple independent hash functions. A Hash Encoder compresses this signature into a single latent vector for processing by a Transformer decoder. Then, a Hash Decoder generates the hash signature of the next token, which is then mapped back to text. We evaluate our approach at the 100M, 1B and 3B parameter scales, demonstrating that MultiHashFormer consistently outperforms standard Transformer LMs across multiple benchmarks. Furthermore, we show that our model handles multilingual vocabulary expansion with a constant parameter footprint without any modifications.
Sources
- Phi-4 Technical Report
- Think you have Solved Direct-Answer Question Answering? Try ARC-DA, the Direct-Answer AI2 Reasoning Challenge
- Olmo 3
- MuBench: Assessment of Multilingual Capabilities of Large Language Models Across 61 Languages
- Dynamic Chunking for End-to-End Hierarchical Sequence Modeling
- Mistral 7B
- Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining
- Bolmo: Byteifying the Next Generation of Language Models
- AdaptiVocab: Enhancing LLM Efficiency in Focused Domains through Lightweight Vocabulary Adaptation
- Continually Adding New Languages to Multilingual Language Models
- FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
- Trans-Tokenization and Cross-lingual Vocabulary Transfers: Language Adaptation of LLMs for Low-Resource NLP
- Qwen3 Technical Report
- From Bytes to Ideas: Language Modeling with Autoregressive U-Nets
- ReCoRD: Bridging the Gap between Human and Machine Commonsense Reading Comprehension
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering