DecompressionLM: Deterministic, Diagnostic, and Zero-Shot Concept Graph Extraction from Language Models
cs.CL
Submitted: 2026-01-30
Updated: 2026-09-12
Code: https://github.com/ulab-uiuc/decompressionlm
License: http://creativecommons.org/licenses/by/4.0/
The gist: Existing knowledge probing methods rely on pre-defined queries, limiting extraction to known concepts.
Terminology
Abstract
Existing knowledge probing methods rely on pre-defined queries, limiting extraction to known concepts. We introduce DecompressionLM, a stateless framework for zero-shot concept graph extraction that discovers what language models encode without pre-specified queries or shared cross-sequence state. Our method targets three limitations of common decoding-based probing approaches: (i) cross-sequence coupling that concentrates probability mass on high-frequency prefixes, (ii) competitive decoding effects that suppress long-tail concepts, and (iii) scalability constraints arising from sequential exploration. Using Van der Corput low-discrepancy sequences with arithmetic decoding, DecompressionLM enables deterministic, embarrassingly parallel generation without shared state across sequences. Across two model families and five quantization variants, we find that activation-aware quantization (AWQ-4bit) expands concept coverage by 30-170%, while uniform quantization (GPTQ-Int4) induces 71-86% coverage collapse - divergent behaviors not reliably reflected by explanation-level perplexity. Corpus-based verification further reveals a 19.6-point hallucination gap between top- and bottom-ranked MMLU-Pro Law models. DecompressionLM establishes concept coverage as a complementary evaluation dimension for assessing knowledge breadth and factual grounding in compressed models intended for deployment. Code is available at https://github.com/ulab-uiuc/decompressionlm.
Sources
- Structured Voronoi Sampling
- TAPAS: Two-pass Approximate Adaptive Sampling for Softmax
- Language Models are Few-Shot Learners
- LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
- QLoRA: Efficient Finetuning of Quantized LLMs
- Sample, Don't Search: Rethinking Test-Time Alignment for Language Models
- The Llama 3 Herd of Models
- The Curious Case of Neural Text Degeneration
- On The Computational Complexity of Self-Attention
- Arithmetic Sampling: Parallel Diverse Decoding for Large Language Models
- Deep sequence models tend to memorize geometrically; it is unclear why
- MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
- How Context Affects Language Models' Factual Predictions
- Decoding-Free Sampling Strategies for LLM Marginalization
- Efficient and Asymptotically Unbiased Constrained Decoding for Large Language Models
- Qwen2.5 Technical Report
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering