Linger and Lose: Knowledge Collapse in Low-Bit Language Models
cs.CL
Submitted: 2026-09-26
Updated: 2026-09-26
Terminology
Sources
- Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws
- Training Dynamics Impact Post-Training Quantization Robustness
- Perplexity Can Miss SAE Feature Damage Under Quantization
- Factual recall in linear associative memories: sharp asymptotics and mechanistic insights
- Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations
- MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
- Spectra: Surprising Effectiveness of Pretraining Ternary Language Models at Scale
- Scaling Laws for Precision
- ParetoQ: Improving Scaling Laws in Extremely Low-bit LLM Quantization
- The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
- How much do language models memorize?
- BitNet b1.58 Reloaded: State-of-the-art Performance Also on Smaller Networks
- (How) Learning Rates Regulate Catastrophic Overtraining
- Mapping the Schedule x Bit-Width Boundary in Sub-100M Quantisation-Aware Training
- Muon Outperforms Adam in Tail-End Associative Memory Learning
- Sharpness-Aware Pretraining Mitigates Catastrophic Forgetting
- Bit-by-Bit: Progressive QAT Strategy with Outlier Channel Splitting for Stable Low-Bit LLMs
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering