TRACES: Proactive Safety Auditing for Multi-Turn LLM Agents via Trajectory-State Modeling
cs.CL, cs.LG
Submitted: 2026-05-26
Updated: 2026-09-19
Comments: EMNLP 2026 Main
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: LLM agents increasingly operate through multi-turn tool use and environment interaction, where safety risks often emerge from intermediate steps long before they surface in the final outcome.
Terminology
Abstract
LLM agents increasingly operate through multi-turn tool use and environment interaction, where safety risks often emerge from intermediate steps long before they surface in the final outcome. Reactive auditing is therefore insufficient: post-hoc diagnosis frequently misses the chance to flag risks while they are unfolding. We propose TRACES, a representation-based proactive auditor that learns prefix-level trajectory risk states from the hidden representations of an observer LLM. TRACES induces latent mechanism features from step representations and models their temporal evolution to estimate whether a partial trajectory is drifting toward unsafe behavior. To sidestep the cost and ambiguity of step-level risk annotation, TRACES is trained with weak trajectory-level supervision while still producing dense prefix-level risk estimates. Across multiple agent safety benchmarks, TRACES improves both full-trajectory safety prediction and proactive risk discrimination. Our analyses further suggest that these risk states can help train a safer agent, highlighting the broader potential of proactive auditing for long-horizon agent safety.
Sources
- AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents
- The Llama 3 Herd of Models
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Why Do Multi-Agent LLM Systems Fail?
- LLMArena: Assessing Capabilities of Large Language Models in Dynamic Multi-Agent Environments
- PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages
- ShieldAgent: Shielding Agents via Verifiable Safety Policy Reasoning
- Probing Latent Subspaces in LLM for AI Security: Identifying and Manipulating Adversarial States
- Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents
- Towards Verifiably Safe Tool Use for LLM Agents
- Not All Language Model Features Are One-Dimensionally Linear
- ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
- The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey
- Beyond Single-Turn: A Survey on Multi-Turn Interactions with Large Language Models
- The Linear Representation Hypothesis and the Geometry of Large Language Models
- Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems
- Reflexion: Language Agents with Verbal Reinforcement Learning
- AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security
- AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents
- AgentBench: Evaluating LLMs as Agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering