Agent Safety From Within: Detecting Harmful Trajectories from LLM Internal States
cs.AI
Submitted: 2026-09-27
Updated: 2026-09-27
Code: https://github.com/CSSLab/Tacit
Terminology
Sources
- Understanding intermediate layers using linear classifier probes
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
- The Internal State of an LLM Knows When It's Lying
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories
- The Llama 3 Herd of Models
- Language Models Represent Space and Time
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- SPIN: Sparsifying and Integrating Internal Neurons in Large Language Models for Text Classification
- MINER: Mining Multimodal Internal Representation for Efficient Retrieval
- ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
- AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security
- AgentAuditor: Human-Level Safety and Security Evaluation for LLM Agents
- AgentLens: Interpretable Safety Steering via Mechanistic Subspaces for Multi-Turn Coding Agent
- Synthetic Persona Pretraining: Alignment from Token Zero
- Steering Language Models With Activation Engineering
- OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety
- SafeGEO: Understanding Generative Engine Optimization Risks in Recommendation Agents
- Qwen3 Technical Report
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection