Dual-Frontier: When Can an Agent Trust Its World Model?
cs.AI
Submitted: 2026-09-22
Updated: 2026-09-24
License: http://creativecommons.org/licenses/by/4.0/
The gist: Learned world models are becoming essential to general-purpose agents: by predicting action consequences, they support planning and decision-making while reducing reliance on costly trial and error.
Terminology
Abstract
Learned world models are becoming essential to general-purpose agents: by predicting action consequences, they support planning and decision-making while reducing reliance on costly trial and error. This reliance creates a fundamental ambiguity: when a world-model-guided decision fails, the trajectory alone may not reveal whether the agent's decision rule or the world model caused the loss. We formalize this failure-attribution problem as a counterfactual decomposition of return loss and prove that its components are not identifiable from passive interaction, even for finite-horizon planners. This obstruction motivates Dual-Frontier, a learning principle that admits a world-model-guided decision only when its predicted advantage exceeds a certified bound on decision-relevant world-model error; otherwise, evidence is allocated to world-model verification. Action-conditioned value bounds and a closed-loop extension guarantee non-decreasing return for admitted decisions. Calibrated gates and simultaneous confidence sequences support adaptive evidence reuse, with sufficient and necessary verification bounds. Controlled learned-model experiments validate the predicted failure modes and certification behavior, while cross-backbone tool-use benchmarks instantiate the same verify-then-promote rule in realistic agent world-model pipelines, consistently improving decision quality and reliability.
Sources
- Policy-Aware Model Learning for Policy Gradient Methods
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- General Agents Contain World Models, even under Partial Observability and Stochasticity
- The Llama 3 Herd of Models
- World Models
- Tools as Continuous Flow for Evolving Agentic Reasoning
- CoMAP: Co-Evolving World Models and Agent Policies for LLM Agents
- Policy and World Modeling Co-Training for Language Agents
- GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving
- Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning
- Qwen3 Technical Report
- Self-Evolving World Models for LLM Agent Planning
- Qwen-AgentWorld: Language World Models for General Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection