Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families
Alizishaan Khatri, Dun Li Chan
cs.LG, cs.AI, cs.CR
Submitted: 2026-08-08
Updated: 2026-08-11
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Khatri et al.
Terminology
Abstract
Khatri et al. (2026) [DOI: 10.1109/DSN-W70714.2026.00027] show that lightweight MLP probes on final-layer activations of a single 8B model (LLaMA-3.1-8B) detect harmful prompts at F1 competitive with guard models 1000x larger, using one probe per benchmark. We reproduce this pipeline end-to-end and extend it along two axes the original study leaves open. First, we test whether the result generalizes across other model architecture and scale by training identical probes on activations from models like Gemma-4-E4B, Mistral-7B-v0.3, and Qwen2-7B, using the three benchmarks (WildJailbreak, BeaverTails, AEGIS 2.0). Second, we test how much of the reported performance is affected by non-determinism during inference by repeating extraction under five random seeds and measuring the variance of F1 scores. Our results reproduce the original LLaMA model benchmarks within 0.37 percentage points of the original F1 scores (and within 0.2 points on BeaverTails). We find that the original MLP probe architecture extends to other model families with F1 scores within a point of the values reported for LLaMA-3.1-8B. Our experiments varying seed values reveal an interesting observation: final token latent vectors remained the same for all tested architectures irrespective of the seed values used.
Sources
- NNsight and NDIF: Democratizing Access to Open-Weight Foundation Model Internals
- DeepContext: Stateful Real-Time Detection of Multi-Turn Adversarial Intent Drift in LLMs
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Probing Latent Subspaces in LLM for AI Security: Identifying and Manipulating Adversarial States
- Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks
- Before the Last Token: Diagnosing Final-Token Safety Probe Failures
- Building Guardrails for Large Language Models
- Harmful Intent as a Geometrically Recoverable Feature of LLM Residual Streams
- ShieldGemma: Generative AI Content Moderation Based on Gemma
- LLMScan: Causal Scan for LLM Misbehavior Detection
- LLMs Encode Harmfulness and Refusal Separately
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks