Compression-Aware Abstention: Teaching LLMs to Refuse When KV-Compression Masks Remove Answer Evidence
cs.CL, cs.LG
Submitted: 2026-08-30
Updated: 2026-08-30
Comments: 19 pages, 5 figures. Accepted to the GroundLM workshop at EMNLP 2026. Code, adapters, and datasets: https://github.com/mali-kh/compression-aware-abstention
Code: https://github.com/mali-kh/compression-aware-abstention
License: http://creativecommons.org/licenses/by/4.0/
The gist: KV-cache compression reduces LLM inference memory by evicting context tokens, but when the evicted tokens contain answer-bearing evidence, the model may hallucinate instead of recognizing that the
Terminology
Abstract
KV-cache compression reduces LLM inference memory by evicting context tokens, but when the evicted tokens contain answer-bearing evidence, the model may hallucinate instead of recognizing that the compressed context is insufficient. We address this failure from a behavioral perspective: to our knowledge, this is the first work to formulate compression-aware abstention as a learning problem, in which a model learns to answer when supporting evidence survives compression and abstain when it does not. We construct supervision from compressor survival masks and tight answer-bearing spans, labeling examples as Confident when evidence survives and Abstain when it is removed. A 10.1M-parameter LoRA adapter trained on 2.6K MuSiQue 2-hop QA examples reduces base-model hallucinations by 97% under prompt-style truncation while preserving correct answering on evidence-retaining examples. Unlike prompt-only abstention baselines, which over-abstain on many answerable high-retention examples, the trained adapter learns a conditional policy. We also evaluate the method under actual compressed-cache decoding, where multi-compressor training yields a 6-22x relative lift over the unaided base on evidence-retaining examples. Controlled-deletion experiments show that the learned behavior is driven by evidence content rather than input length alone.
Sources
- Knowledge of Knowledge: Exploring Known-Unknowns Uncertainty with Large Language Models
- Sufficient Context: A New Lens on Retrieval Augmented Generation Systems
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
- Language Models (Mostly) Know What They Know
- PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
- KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction
- KV-Distill: Nearly Lossless Learnable Context Compression for LLMs
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- SnapKV: LLM Knows What You are Looking for Before Generation
- Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution
- Teaching Models to Express Their Uncertainty in Words
- Scissorhands: Exploiting the Persistence of Importance Hypothesis for LLM KV Cache Compression at Test Time
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
- LoRA: Low-Rank Adaptation of Large Language Models
- The Llama 3 Herd of Models
- Qwen2.5 Technical Report
- Self-Evaluation Improves Selective Generation in Large Language Models
- Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering