LLM Safety From Within: Detecting Harmful Content with Internal Representations
cs.AI
Submitted: 2026-04-20
Updated: 2026-04-20
Comments: 17 pages,10 figures,6 tables
Journal ref: Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2026
DOI: 10.18653/v1/2026.acl-long.1844
Code: https://github.com/CSSLab/SIREN
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Guard models are widely used to detect harmful content in user prompts and LLM responses.
Terminology
Abstract
Guard models are widely used to detect harmful content in user prompts and LLM responses. However, state-of-the-art guard models rely solely on terminal-layer representations and overlook the rich safety-relevant features distributed across internal layers. We present SIREN, a lightweight guard model that harnesses these internal features. By identifying safety neurons via linear probing and combining them through an adaptive layer-weighted strategy, SIREN builds a harmfulness detector from LLM internals without modifying the underlying model. Our comprehensive evaluation shows that SIREN substantially outperforms state-of-the-art open-source guard models across multiple benchmarks while using 250 times fewer trainable parameters. Moreover, SIREN exhibits superior generalization to unseen benchmarks, naturally enables real-time streaming detection, and significantly improves inference efficiency compared to generative guard models. Overall, our results highlight LLM internal states as a promising foundation for practical, high-performance harmfulness detection.
Sources
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts
- Aegis2.0: A Diverse AI Safety Dataset and Risks Taxonomy for Alignment of LLM Guardrails
- ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
- Do Internal Layers of LLMs Reveal Patterns for Jailbreak Detection?
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Linearity of Relation Decoding in Transformer Language Models
- PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages
- Beyond the Surface: Enhancing LLM-as-a-Judge Alignment with Human via Internal Representations
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
- Safety Layers in Aligned Large Language Models: The Key to LLM Security
- ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- The Linear Representation Hypothesis and the Geometry of Large Language Models
- Lightweight Safety Classification Using Pruned Language Models
- Layer by Layer: Uncovering Hidden Representations in Language Models
- Linear Representations of Sentiment in Large Language Models
- Qwen3 Technical Report
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection