Local Sparsity Enables Unsupervised LLM Safety Detection
cs.LG, cs.AI
Submitted: 2026-09-17
Updated: 2026-09-17
Code: https://github.com/lasgroup/unsupervised-llm-safety
License: http://creativecommons.org/licenses/by/4.0/
The gist: Deployment-time safety methods for large language models (LLMs) are predominantly supervised and assume access to unsafe training data.
Terminology
Abstract
Deployment-time safety methods for large language models (LLMs) are predominantly supervised and assume access to unsafe training data. Nevertheless, new attacks and harm categories regularly arise, not captured by models trained in such a supervised fashion. An alternative approach is to view this problem through the lens of anomaly detection, namely, to rely solely on modeling safe data and flagging out-of-distribution inputs. However, LLM activations lie in a high-dimensional space, raising concerns about whether anomaly detection is statistically feasible. We show that, under the linear representation hypothesis (LRH), there may indeed be hope. In the LRH concept space, which is typically recovered via a sparse autoencoder (SAE), nearby points share a small common active support. Using this local sparsity insight, we propose a framework for locally masked SAE-based anomaly detection, supported by theoretical justifications. We validate it on various architectures and datasets, including both capability-testing datasets and safety-specific datasets. Finally, when we allow algorithms to use 1% out-of-distribution data for calibration, locally sparse methods achieve near-optimal performance, demonstrating their ability to capture meaningful safety information while using only 1-2% of SAE neurons for computation.
Sources
- gpt-oss-120b & gpt-oss-20b Model Card
- Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
- Program Synthesis with Large Language Models
- Qwen Technical Report
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
- AI Deception: Risks, Dynamics, and Controls
- Evaluating Large Language Models Trained on Code
- Learning to Inject: Automated Prompt Injection via Reinforcement Learning
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Safe RLHF: Safe Reinforcement Learning from Human Feedback
- Toy Models of Superposition
- Scaling and evaluating sparse autoencoders
- The Llama 3 Herd of Models
- Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
- Measuring Massive Multitask Language Understanding
- Out-of-distribution Detection in Medical Image Analysis: A survey
- Specialization after Generalization: Towards Understanding Test-Time Training in Foundation Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks