Language Models Without a Trainable Input Embedding Table: Learning from Fixed Minimal Binary Token Codes
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Language Models Without a Trainable Input Embedding Table".
Jane: Trainable input embedding tables are standard in language models,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: We’re starting by looking at the title and who wrote this paper, "Language Models Without a Trainable Input Embedding Table: Learning from Fixed Minimal Binary Token Codes." It immediately signals that the core focus is on removing that trainable input embedding table from the system.
Jane: And you see how they frame it around using fixed minimal binary token codes instead of learned vectors, which is a very direct way of pointing out their central research question.
Lu: The authors are tackling a long-standing convention in the field where token IDs are mapped through a learned input table before any contextual processing happens, and this paper challenges that necessity.
Meng: I see them focusing on V = sixty-five thousand five hundred thirty-six for their main test setting and dmodel = one thousand twenty-four which gives us some concrete numbers to work with when thinking about the scale of this experiment <ref:2605.09751#pg0>.
Lalam: It’s fascinating that they are comparing this fixed code approach against a standard learned-input baseline to see if it can achieve comparable performance without any trainable input parameters whatsoever.
The paper's summary: Tom: To summarize what the paper is actually doing, they propose replacing the usual trainable V times dmodel input embedding matrix with fixed minimal binary token codes and a zero-parameter lift to model width.
Jane: Essentially, instead of learning a continuous vector for each token ID, they use a deterministic way to represent tokens using only the exact identity information.
Lu: They define this by noting that for any vocabulary size V, exact token identity requires only K = two V bits, and they use the binary expansion of the token ID as their minimal code <ref:2605.09751#pg0,size $V$, exact token identity requires only $K>.
Meng: The lift construction they describe is very specific, using a deterministic tiling where the K-bit code is repeated until it reaches the model width of one thousand twenty-four <ref:2605.09751#pg0>. That ensures every input vector stays within a subspace of dimension at most K.
Lalam: This means even though the model has a large width, all its inputs are constrained to this very small dimensional space defined by the token identity itself.
The paper's improvements: Tom: The main improvement they highlight is that their fixed-code variants are competitive with the standard baseline without introducing any trainable input parameters at all.
Jane: They found that in their experiments, the fixed-code model actually had a lower mean validation perplexity of two point three six compared to two point four four for the learned-input baseline, which is a small difference across three independent seeds <ref:2605.09751#pg0,a lower mean validation perplexity>.
Lu: That result strongly supports the idea that exact token identity delivered through these fixed minimal binary codes is sufficient for the Transformer stack to learn effective internal representations, even without those initial learned vectors.
Meng: The paper also presents a fully table-free variant where the codes are generated on the fly and optionally recoded by an invertible affine transform over F K squared, which removes any special dependence on the canonical binary ordering of token IDs <ref:2605.09751#pg0,recoded by an invertible affine transform over $F_K^2>.
Lalam: That ability to test robustness against code assignment via that affine recoding experiment is a really smart addition because it shows performance isn't tied to a specific, perhaps arbitrary, way the codes are ordered.
Conclusion: Tom: So, wrapping up on this paper about "Language Models Without a Trainable Input Embedding Table: Learning from Fixed Minimal Binary Token Codes," the big implication is that for certain setups, we don't need those large trainable input embedding tables to get useful language modeling results.
Jane: It really suggests that the model learns representations from a stable symbolic interface where tokens enter only as an injective minimal binary code, which simplifies how we view input representation.
Lu: This provides a constructive counterexample showing that in this specific regime, a trainable input embedding table is not necessary for achieving good performance on downstream tasks.
Meng: Practically, it means we can remove sixty-seven point one million trainable input parameters and save significant memory and training time while maintaining quality in this setting.
Lalam: For our culture, this points toward building models that are inherently simpler at the input level, focusing the learning power entirely on the complex contextual processing within the Transformer stack itself.
Tom: It's a lot to take in today about how fundamental input structure can be simplified. We’ll keep an eye on how this idea plays out with other architectures next time.
cs.CL
Submitted: 2026-05-10
Updated: 2026-10-02
Comments: Updated version: added standardized base-model benchmark evaluations, expanded related work, corrected references, and clarified parameter accounting and experimental limitations
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 89/100
The gist: Trainable input embedding tables are standard in language models, but this work investigates whether they are necessary by replacing them with fixed minimal binary token codes and finding that exact
Key concepts
- Learned Input Table
- This is the standard method where every unique token ID is mapped to a specific learned vector in a large lookup table. It requires training these vectors, which adds millions of trainable parameters to the model's input layer.
- Fixed Minimal Binary Code
- Instead of learning embeddings, this method uses the binary expansion of a token ID (e.g., 1024 tokens require 10 bits). This fixed code is then deterministically tiled and lifted to match the model's dimension, ensuring exact token identity is preserved without any trainable parameters.
- Affine-Recoded Table-Free Code
- This advanced variant computes the minimal binary code on the fly and optionally applies an invertible linear transformation over F2. This approach is fully table-free, meaning it avoids storing any precomputed input lookups while still maintaining robustness against different code assignments.
Terminology
Summary
Trainable input embedding tables are standard in language models, but this work investigates whether they are necessary by replacing them with fixed minimal binary token codes and finding that exact token identity is sufficient for useful language modeling.
The gist: In matched 32-layer experiments with V = 65,536, fixed 16-dimensional binary codes tiled to dmodel = 1024 match the learned-input baseline while eliminating 67.1M trainable input parameters.
How it works
The core idea is to replace the standard trainable input lookup, which maps token IDs to a learned vector, with a deterministic binary coding interface that preserves exact token identity. For a vocabulary of size V, this requires only K = ⌈log2 V⌉ bits for exact identity. The authors use the canonical minimal code where the K-bit code is the binary expansion of the token ID: c(t) = binK(t) ∈ 0, 1K
. This ensures that exact token identity requires only K = ⌈log2 V⌉ bits.
The resulting input representation, x(t), is constructed by applying a fixed zero-parameter lift to this binary code. This lift maps the K-dimensional code space into the model width dmodel. The authors specify a deterministic tiled lift: R = [IK IK... IK] ∈ R d×K,
where IK is the K × K identity matrix repeated s times vertically, ensuring that the K-bit code is repeated until the model width is reached.
This construction guarantees that all input vectors lie in a subspace of dimension at most K, even though the model width is much larger.
Key Implementations and Variants
The study evaluates three primary input parameterizations. First, the standard Learned input table,
which has a size of 65,536 × 1024 = 67.1M trainable input parameters.
Second, the Fixed minimal binary code,
which replaces this with zero trainable parameters by using the deterministic tiling described above. Third is the Affine-recoded table-free code,
which computes codes on the fly and optionally applies an invertible affine transform over F2: c˜(t) = A c(t) ⊕ b, A ∈ GL(K, 2), b ∈ 0, 1K.
This variant is fully table-free.
Results and Findings
The main empirical finding is that the fixed-code variants are competitive with the standard baseline without any trainable input parameters. Across three independent seeds, the fixed-code model has a lower mean validation perplexity in our runs, 2.36 versus 2.44 for the learned-input baseline,
but this difference is within the measured seed-to-seed variation of 4.8%.
The affine-recoded table-free variant also performs well, achieving validation perplexity 2.39 with zero trainable input parameters.
This demonstrates that exact token identity delivered through fixed minimal binary codes is sufficient for the downstream Transformer stack to learn effective internal representations.
Robustness and Interpretation
The authors test robustness against code assignment via the affine-recoded experiment. They found that the affine-recoded experiment provides a direct robustness check on the code assignment itself,
observing no collapse
when using an invertible transform over F162, suggesting that performance is not dependent on a privileged token-ID ordering or to a hand-designed semantic geometry at the input.
The paper concludes that this result serves as a constructive counterexample to necessity,
showing that in this regime, a trainable input embedding table is not necessary for useful language modeling.
The key takeaway is that the model learns representations from a stable symbolic interface
where tokens enter only as an injective minimal binary code.
Limitations and Broader Impact
Limitations include the reliance on a single tokenizer family and vocabulary size, and the lack of extensive hyperparameter retuning for each input parameterization. However, the work has broader impact by providing a simpler and more controlled way to study what information must be present at the input of a language model,
potentially reducing optimizer-state requirements associated with the input layer.
The method does not introduce new capabilities but rather clarifies that token representation at the input need not be a semantic vector, but can simply be a stable address. The results show that non-degradation relative to a learned-input baseline is itself the central result,
supporting the claim that the standard trainable input table can be removed without degrading held-out next-token modeling quality in this setting.
The output projection remains standard and trainable throughout. No new personally identifying datasets or data collection procedures were introduced. The work provides a constructive counterexample to necessity
rather than a benchmark optimization claim, focusing on the existence of a working model without a trainable input lookup.
Improvements for AI systems
Based on the scientific paper, here are specific improvements that can be made to AI systems by implementing the proposed input parameterization:
-
Improve Memory Efficiency and Training Speed: By replacing a massive trainable input embedding matrix (67.1M parameters for V=65,536, d=1024) with a fixed, zero-parameter binary code interface, the system saves 67.1 million trainable parameters and associated memory overhead in the input layer. This leads to faster training convergence and lower hardware requirements for storing model weights during pretraining.
-
Enhance Robustness to Input Noise: Since the input representation is derived from an exact, minimal binary code (or an affinely recoded version thereof), it removes the dependency on learning a continuous
semantic vector
from noisy or fragmented token IDs (like unusual casing or spelling variations). The model's ability to learn meaning relies solely on contextual computation within the Transformer stack, making the input layer more robust to lexical ambiguity and minor tokenization noise. -
Improve Interpretability of Input Structure: The system gains a clear understanding that for exact token identity, only a minimal set of bits (K=16 for V=65,536) is required at the input interface. This isolates the necessity of learned continuous vectors from the necessity of stable symbolic addresses, clarifying whether meaning is constructed through contextual computation rather than direct retrieval from a learned lookup table.
-
Enable Table-Free/On-the-Fly Inference: The model can be deployed with zero trainable input parameters, allowing for highly efficient inference where the input representation is computed directly from the token ID at runtime using the fixed lift matrix and binary code extraction logic. This is particularly valuable for edge devices or low-latency applications where loading large embedding tables is prohibitive.
-
Improve Transferability Across Code Assignments: By testing an invertible affine recoding over F2, the system demonstrates that performance is not critically dependent on the canonical ordering of token IDs or a specific geometric structure of the binary codebook. This suggests that researchers can use randomly assigned or adversarially perturbed code assignments without expecting catastrophic performance degradation, increasing flexibility in input representation design.
-
Support Adaptation Regimes Effectively: The fixed-code interface proves that a model can be trained from scratch with a minimal input parameterization and then successfully adapted via Supervised Fine-Tuning (SFT) to downstream tasks (as shown in Table 5). This provides a stable, minimal starting point for transfer learning without the risk of overfitting or instability introduced by a large, complex learned embedding table.
Abstract
We study whether a decoder-only language model requires an independently trainable input vector for every token. For a vocabulary of size V, an injective fixed-length binary identifier requires K= 2 V bits. We replace the usual trainable V times d model input table with fixed minimal binary token codes and a parameter-free tiled lift to model width. With V=65, 536 and d model=1024, this supplies each token as a fixed 16-bit code and removes 67.1M trainable parameters, approximately 12.5% of the untied learned-input baseline. We also study a table-free implementation with one fixed invertible affine recoding over F 2 16. Across three training seeds, 32-layer models trained on approximately 16-17B tokens obtain mean held-out perplexities of 2.44 for the learned-input baseline, 2.36 for canonical binary codes, and 2.39 for affine-recoded codes. These descriptive results do not establish statistical superiority or equivalence. Standardized evaluation of released base checkpoints with the LM Evaluation Harness adds commonsense, knowledge, and language-modeling benchmarks. The three paper checkpoints show broadly similar, task-dependent performance, while external SmolLM2 reference models are substantially stronger on many tasks. Our conclusion is therefore limited to the studied regime: a free trainable token-indexed input table is not required to learn nontrivial language modeling. The Transformer still learns continuous representations, and the output vocabulary projection remains standard and trainable.
Sources
- Efficient Estimation of Word Representations in Vector Space
- The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
- MEGABYTE: Predicting Million-byte Sequences with Multiscale Transformers
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering