The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators
cs.CR, cs.AI, cs.LG
Submitted: 2026-07-12
Updated: 2026-09-27
Comments: 19 pages, 4 figures, 30 tables. Code and supporting artifacts: https://github.com/dschwarz32/entanglement-wall
Code: https://github.com/dschwarz32/entanglement-wall
License: http://creativecommons.org/licenses/by/4.0/
The gist: Context can change whether a request is harmful without changing its topic or surface form.
Terminology
Abstract
Context can change whether a request is harmful without changing its topic or surface form. We ask whether residual-stream probes distinguish harmful requests from surface-matched benign controls at a useful operating point. Across three 7-8B model families, an activation sensor blocks 95.5-97.7 percent of judge-classified compliant attacks in a taxonomy-selected set. It also blocks 59.6-68.4 percent of XSTest prompts. A fully disjoint audit reconstructs near-ceiling source-contrast AUROC (0.996-0.999), but fixed transfer to matched pairs is weaker: 0.656-0.819 on the guard-selected Twin-n70 subset and 0.590-0.690 on the full Twin-n163 cohort. We test ten axes on the reference family and seven across all families with leakage, hold-out, and permutation controls. On Twin-n163, no axis evaluated without direct pair-boundary fitting reaches the specified numerical threshold. Requiring persistence on that full cohort was added at analysis time. A separately specified 24B/32B extension gives the same result. Pair-trained classifiers weaken under category and generation-batch hold-out and false-block 79.6-100 percent of XSTest at 95 percent in-corpus TPR. At the tested read points, these activation scores behave as broad-risk detectors rather than standalone context adjudicators.
Sources
- Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
- Refusal in Language Models Is Mediated by a Single Direction
- An Interpretability Illusion for BERT
- With Little Power Comes Great Responsibility
- JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
- Defeating Prompt Injections by Design
- When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift
- Evaluating Models' Local Decision Boundaries via Contrast Sets
- WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
- Segment-Level Coherence for Robust Harmful Intent Probing in LLMs
- Mistral 7B
- WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
- There Is More to Refusal in Large Language Models than a Single Direction
- Learning the Difference that Makes a Difference with Counterfactually-Augmented Data
- Building Production-Ready Probes For Gemini
- Probing Classifiers are Unreliable for Concept Removal and Detection
- On Calibration of LLM-based Guard Models for Reliable Content Moderation
- Harmful Intent as a Geometrically Recoverable Feature of LLM Residual Streams
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- Detecting High-Stakes Interactions with Activation Probes
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs