The Role of Fine-grained Harm Signals in LLM Safety
cs.CL
Submitted: 2026-09-16
Updated: 2026-09-18
Comments: 9 pages, 6 figures
Code: https://github.com/tatsu-lab/stanford_alpaca
License: http://creativecommons.org/licenses/by/4.0/
The gist: Prior work has shown that internal harmfulness representations in large language models vary across risk categories, while sharing a common general harm representation component.
Terminology
Abstract
Prior work has shown that internal harmfulness representations in large language models vary across risk categories, while sharing a common general harm representation component. This raises a question about the role of the category-specific component beyond general harm representation in LLM safety. To answer this question, we isolate the category-specific component by removing shared general harmfulness representation from each categorical harmfulness representation, yielding a category residual that is orthogonal to general harmfulness at every layer. Using activation steering with category residuals across 11 risk categories in 3 instruction-tuned LLMs, we find that whether category residuals encode harmfulness varies across categories, and that this category-wise pattern is similar across models. Whether category residuals induce refusal also varies across categories, but this category-wise pattern is more model-dependent. We also find that category residuals increase LLMs' downstream internal alignment with shared general harmfulness representation. Together, these findings demonstrate that more fine-grained category residuals should also be considered beyond shared general harmfulness representation to fully understand LLM safety. More broadly, our findings show that even a direction orthogonal to a concept at one layer can contribute to the concept's downstream amplification.
Sources
- Gemma 2: Improving Open Language Models at a Practical Size
- The Llama 3 Herd of Models
- The Geometry of Harmfulness in LLMs through Subconcept Probing
- Qwen2.5 Technical Report
- LLMs Encode Harmfulness and Refusal Separately
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering