Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models
cs.AI, cs.CL, cs.CR, cs.LG
Submitted: 2026-09-16
Updated: 2026-09-16
Journal ref: IEEE DSN-W 2026, pp. 48-52
DOI: 10.1109/DSN-W70714.2026.00027
License: http://creativecommons.org/licenses/by/4.0/
The gist: Autonomous systems increasingly rely on Large Language Models (LLMs) yet the safety infrastructure surrounding these models introduces latency and compute overhead.
Terminology
Abstract
Autonomous systems increasingly rely on Large Language Models (LLMs) yet the safety infrastructure surrounding these models introduces latency and compute overhead. This limits utility in resource-constrained, time-critical deployments. Existing external guardrail models remain blind to the model's internal workings, creating a fundamental assurance gap. We ask: does the model already know when the content is harmful? We extract activations from LLaMA-3.1-8B and train lightweight MLP classifier probes (12.6M parameters) to detect harmful prompts. Evaluated on WildJailbreak, Beavertails, and AEGIS 2.0, our probes achieve F1 scores of 99%, 83%, and 84%, respectively competitive with 1000x larger guard models while cutting latency and compute costs.
Sources
- Next-Generation LLM for UAV: From Natural Language to Autonomous Flight
- Jailbreaking Attack against Multimodal Large Language Model
- Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Alignment faking in large language models
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- ShieldGemma: Generative AI Content Moderation Based on Gemma
- Building Guardrails for Large Language Models
- Representation Engineering: A Top-Down Approach to AI Transparency
- LLMScan: Causal Scan for LLM Misbehavior Detection
- Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks
- DeepContext: Stateful Real-Time Detection of Multi-Turn Adversarial Intent Drift in LLMs
- Lightweight Safety Guardrails Using Fine-tuned BERT Embeddings
- Llama Guard 3-1B-INT4: Compact and Efficient Safeguard for Human-AI Conversations
- Hiding in Plain Sight: A Steganographic Approach to Stealthy LLM Jailbreaks
- Bypassing Prompt Guards in Production with Controlled-Release Prompting
- AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
- Probing Latent Subspaces in LLM for AI Security: Identifying and Manipulating Adversarial States
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection