JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety
Yuan Xiong, Linji Hao, Shizhu He, Yequan Wang, Lijun Li
cs.AI, cs.CL, cs.CR
Submitted: 2026-07-22
Code: https://github.com/xiongyuaay/JANUS
License: http://creativecommons.org/licenses/by/4.0/
The gist: Agent safety is moving from content moderation toward preventing operational failures before tool-using agents act.
Terminology
Abstract
Agent safety is moving from content moderation toward preventing operational failures before tool-using agents act. We propose Janus, a foresight-oriented framework for long-horizon agent safety that trains guards to anticipate delayed risks from partial trajectories. Janus synthesizes diverse agent trajectories via multi-agent simulation and learns a shared policy with two coupled tasks: an anticipation task that forecasts safety-relevant futures and an adjudication task that decides safety from both the observed prefix and anticipated future. The two tasks are jointly optimized with CoAA-RL, which rewards forecasts by their utility for downstream safety judgment. The resulting guard model, Vanguard, blocks unsafe actions before execution. Across four agent-safety benchmarks, Vanguard improves average protection by 15.9 percentage points over baseline guards while increasing benign task completion by 5.1 percentage points.
Sources
- Constitutional AI: Harmlessness from AI Feedback
- Evaluating Large Language Models Trained on Code
- LPS-Bench: Benchmarking Safety Awareness of Computer-Use Agents in Long-Horizon Planning under Benign and Adversarial Scenarios
- SafePred: A Predictive Guardrail for Computer-Using Agents via World Models
- ShieldAgent: Shielding Agents via Verifiable Safety Policy Reasoning
- LlamaFirewall: An open source guardrail system for building secure AI agents
- The Llama 3 Herd of Models
- Building a Foundational Guardrail for General Agentic Systems via Synthetic Data
- AgentLAB: Benchmarking LLM Agents against Long-Horizon Attacks
- AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security
- Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- ToolSafe: Enhancing Tool Invocation Safety of LLM-based agents via Proactive Step-level Guardrail and Feedback
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- TabSieve: Explicit In-Table Evidence Selection for Tabular Prediction
- From Assistant to Double Agent: Formalizing and Benchmarking Attacks on OpenClaw for Personalized Local AI Agent
- GuardAgent: Safeguard LLM Agents by a Guard Agent via Knowledge-Enabled Reasoning
- Stable Adaptive Thinking via Advantage Shaping and Length-Aware Gradient Regulation
- Qwen3 Technical Report
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection