Baseline Shape Decides the Verdict: A Controlled Re-Examination of Ternary Language Models at 60K Parameters
cs.CL
Submitted: 2026-09-24
Updated: 2026-09-24
Code: https://github.com/veldanda/ByteLM
Terminology
Sources
- Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
- Hymba: A Hybrid-head Architecture for Small Language Models
- TinyStories: How Small Can Language Models Be and Still Speak Coherent English?
- Ternary Mamba: Grouped Quantization-Aware Training of W1.58A16 State Space Models
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- Training Compute-Optimal Large Language Models
- Spectra: Surprising Effectiveness of Pretraining Ternary Language Models at Scale
- ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization
- The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
- TernaryLM: Memory-Efficient Language Modeling via Native 1.5-Bit Quantization with Adaptive Layer-wise Scaling
- Scaling Laws vs Model Architectures: How does Inductive Bias Influence Scaling?
- Mapping the Schedule x Bit-Width Boundary in Sub-100M Quantisation-Aware Training
- Spectra 1.1: Scaling Laws and Efficient Inference for Ternary Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering