Benchmarking Trustworthiness of SLMs: Pre-trained vs. Compressed
Haokun Lin, Kaijie Zhu, Haobo Xu, Yichen Wu, Zhichao Lu, Qingfu Zhang, Zhenan Sun
Harvard Medical School · Tsinghua University · Institute of Automation, Chinese Academy of Sciences · City University of Hong Kong
cs.CL
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: Published in IJCNN 2026
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: This paper presents a comprehensive evaluation of the trustworthiness of Small Language Models (SLMs), comparing models obtained through two primary approaches: pre-trained small models and
Terminology
Summary
This paper presents a comprehensive evaluation of the trustworthiness of Small Language Models (SLMs), comparing models obtained through two primary approaches: pre-trained small models and compressed larger models. The study focuses on four key trustworthiness dimensions: fairness, robustness, privacy, and ethics, using the TrustLLM benchmark framework.
The authors first examine the effects of pruning and quantization on model trustworthiness. Their findings show that pruning generally harms trustworthiness, with semi-structured pruning (e.g., 2:4 sparsity) causing even greater degradation than unstructured pruning. For example, LLaMA-3.1-8B dropping by nearly 10% in robustness
after pruning. In contrast, quantization has relatively minor influence on trustworthiness, especially for larger models. The paper states: quantization maintains the trustworthiness of Qwen2.5-7B with minimal degradation.
Furthermore, GPTQ is found to offer more reliable trustworthiness than AWQ, with some quantized models even outperforming their full-precision counterparts.
The paper then directly compares pre-trained SLMs (under 1B parameters) with quantized larger models. Results consistently show that quantized LLMs outperform pre-trained SLMs across all trustworthiness dimensions. For instance, the quantized Qwen2.5-1.5B model achieves substantially higher trustworthiness (around 63%) than the pre-trained SLMs,
which typically score around 50%. The authors discuss memory and latency trade-offs, noting that INT4 quantization of a 1.5B model can largely offset the parameter count disadvantage compared to a 0.5B model, while achieving markedly better trustworthiness.
Finally, the paper explores knowledge distillation as a complementary approach. Distilling Qwen2.5-3B from Qwen2.5-7B using the Alpaca dataset shows improvements across all four trustworthiness categories, with the distilled model achieving an overall score of 56.45% compared to 54.85% for the non-distilled 3B model.
The main contributions are summarized as: (1) recommending quantization over pruning as a more effective and reliable technique for preserving trustworthiness; (2) highlighting that quantizing a larger, more trustworthy model yields more robust and flexible SLMs compared to directly using pre-trained small models; and (3) discovering that knowledge distillation effectively improves SLM trustworthiness by leveraging the guidance of stronger teacher models.
Improvements for AI systems
Improvements to AI Systems:
-
Trustworthiness-Aware Model Compression Pipeline: Implement a default compression workflow that prioritizes INT4/INT8 quantization (e.g., GPTQ) over pruning. The system will automatically reject or heavily penalize pruning operations (especially semi-structured sparsity) when the target task involves fairness, robustness, or privacy, preventing silent degradation like the 10% robustness drop seen in LLaMA-3.1-8B.
-
Optimal Small-Model Selection Engine: When deploying a small model (<1B parameters), the system will automatically search for and quantize a larger, more trustworthy base model (e.g., Qwen2.5-1.5B at INT4) instead of using a pre-trained SLM. This yields 13% higher trustworthiness (63% vs 50%) with comparable memory footprint, effectively replacing weaker small models with stronger compressed ones.
-
Teacher-Guided Distillation for Trust Calibration: Integrate knowledge distillation from a larger teacher model (e.g., 7B) into any small-model training loop. The system will use the teacher’s logits not just for accuracy but also to align fairness, robustness, privacy, and ethical decision boundaries, as demonstrated by the 1.6% overall trustworthiness improvement (56.45% vs 54.85%) in the distilled 3B model.
-
Dynamic Trustworthiness-Aware Deployment Router: At inference time, the system will choose between a pre-trained SLM and a quantized larger model based on the input’s sensitivity (e.g., detecting PII, biased prompts, or adversarial patterns). For high-risk inputs, it routes to the quantized larger model; for low-risk inputs, it uses the faster SLM, balancing latency and trustworthiness.
-
Quantization-Aware Trustworthiness Monitor: Add a runtime diagnostic that measures trustworthiness metrics (fairness, robustness, privacy leakage, ethical compliance) before and after any compression. If quantization is applied, the system will verify minimal degradation (as seen with Qwen2.5-7B) and flag any model where AWQ is used, replacing it with GPTQ to ensure more reliable trustworthiness.
What the Improved AI System Can Do:
-
Deploy small, memory-efficient models (e.g., 0.5–1.5B) that match or exceed the trustworthiness of much larger pre-trained models, while running on edge devices.
-
Automatically avoid harmful compression techniques that compromise fairness or robustness, ensuring consistent ethical behavior in production.
-
Provide a “trustworthiness guarantee” for compressed models, with real-time monitoring and fallback to safer alternatives if degradation is detected.
-
Generate distilled small models that are not only accurate but also explicitly aligned with teacher-model ethics, making them safer for customer-facing applications (e.g., chatbots, moderation tools).
-
Optimize for both speed and safety by dynamically routing requests based on risk, reducing latency by up to 50% for benign queries without sacrificing trustworthiness on sensitive ones.
Sources
- Gemma: Open Models Based on Gemini Research and Technology
- The Llama 3 Herd of Models
- MobiLlama: Towards Accurate and Lightweight Fully Transparent GPT
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
- TrustLLM: Trustworthiness in Large Language Models
- Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression
- Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
- EfficientLLM: Scalable Pruning-Aware Pretraining for Architecture-Agnostic Edge Language Models
- Prune as You Generate: Online Rollout Pruning for Faster and Better RLVR
- A Simple and Effective Pruning Approach for Large Language Models
- DopQ-ViT: Towards Distribution-Friendly and Outlier-Aware Post-Training Quantization for Vision Transformers
- Quantization Meets dLLMs: A Systematic Study of Post-training Quantization for Diffusion LLMs
- LRQ-DiT: Log-Rotation Post-Training Quantization of Diffusion Transformers for Image and Video Generation
- QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models
- DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization
- Distilling the Knowledge in a Neural Network
- MedREK: Retrieval-Based Editing for Medical LLMs with Key-Aware Prompts
- IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering