Debias-SparseGPT: Bias-Aware Pruning for Large Language Models
cs.CL
Submitted: 2026-09-02
Updated: 2026-09-02
Comments: Accepted to EMNLP 2026, Code: https://github.com/upunaprosk/debias-llm-compressor
Code: https://github.com/upunaprosk/debias-llm-compressor
License: http://creativecommons.org/licenses/by/4.0/
The gist: Model compression techniques such as pruning and quantization facilitate the efficient deployment and acceleration of Large Language Models (LLMs).
Terminology
Abstract
Model compression techniques such as pruning and quantization facilitate the efficient deployment and acceleration of Large Language Models (LLMs). However, recent studies show that weight sparsification methods, such as SparseGPT, can amplify existing biases in models, with outputs varying significantly depending on persona cues in the prompt. In this paper, we introduce Debias-SparseGPT, a post-training pruning method incorporating representational debiasing using a second-order term defined over demographically contrasting inputs. We perform empirical validation of our method over a wide range of generative LLMs. Across models and sparsity regimes (25%, 50%, and structured 2:4 sparsity), Debias-SparseGPT consistently reduces pruning-induced bias compared to SparseGPT while preserving model perplexity and zero-shot accuracy. Under the most restrictive 2:4 structured sparsity pattern, which most aggressively degrades model quality, augmenting the calibration set with long-context, content-rich examples further improves both downstream performance and fairness. Overall, Debias-SparseGPT advances the bias-performance trade-off while preserving the computational efficiency of sparse models.
Sources
- Characterising Bias in Compressed Models
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- GPT-4 Technical Report
- To prune, or not to prune: exploring the efficacy of pruning for model compression
- Can Model Compression Improve NLP Fairness
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering