HINTBench: Horizon-agent Intrinsic Non-attack Trajectory Benchmark
cs.LG, cs.AI
Submitted: 2026-04-15
Updated: 2026-09-16
License: http://creativecommons.org/licenses/by/4.0/
The gist: Existing agent-safety evaluation has focused mainly on externally induced risks.
Terminology
Abstract
Existing agent-safety evaluation has focused mainly on externally induced risks. Yet agents may still enter unsafe trajectories under benign conditions. We study this complementary but underexplored setting through the lens of intrinsic risk, where intrinsic failures remain latent, propagate across long-horizon execution, and eventually lead to high-consequence outcomes. To evaluate this setting, we introduce non-attack intrinsic risk auditing, a guard-oriented safety evaluation task, and present HINTBench, a benchmark of 596 agent trajectories, comprising 400 synthetic risky trajectories, 136 synthetic safe trajectories, 30 reconstructed real-world risky trajectories, and 30 reconstructed real-world safe trajectories, with an average length of 24.0 steps. HINTBench supports three tasks: risk detection, risk-step localization, and intrinsic failure-type identification, with annotations organized under a unified five-constraint taxonomy. Experiments reveal a substantial capability gap: strong LLMs perform well on trajectory-level risk detection, but the best model remains below 37 on fine-grained Strict-F1 for risk-step localization. Existing off-the-shelf guard models evaluated under their native prompts transfer poorly to this setting. These findings establish intrinsic risk auditing as an open challenge for agent safety.
Sources
- AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
- Why Do Multi-Agent LLM Systems Fail?
- ShieldAgent: Shielding Agents via Verifiable Safety Policy Reasoning
- Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding Conversations
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
- GLM-5: from Vibe Coding to Agentic Engineering
- The Llama 3 Herd of Models
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Mistral 7B
- AgentLAB: Benchmarking LLM Agents against Long-Horizon Attacks
- PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages
- AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments?
- A Benchmark for Evaluating Outcome-Driven Constraint Violations in Autonomous AI Agents
- Evaluating Test-Time Scaling of General LLM Agents
- AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security
- ERNIE 5.0 Technical Report
- TrajAD: Trajectory Anomaly Detection for Trustworthy LLM Agents
- AgentAuditor: Human-Level Safety and Security Evaluation for LLM Agents
- AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks