How Strongly Should Task State Influence an LLM Agent?
cs.AI, cs.CL, cs.LG
Submitted: 2026-09-22
Updated: 2026-09-22
Comments: Preprint. 43 pages
Code: https://github.com/amazon-agi/tau2-bench-verified
License: http://creativecommons.org/licenses/by/4.0/
The gist: Long-horizon assigned work requires an LLM agent to track the state of a task: which steps are done, blocked, cancelled, or open to repetition.
Terminology
Abstract
Long-horizon assigned work requires an LLM agent to track the state of a task: which steps are done, blocked, cancelled, or open to repetition. Agent systems either keep this state as text in the prompt and rely on the model to read that text, or move the state into a module that enforces it, and each system is evaluated as a whole, so no one knows how much reliability comes from the state being shown, told, or enforced. We fix the task rules, the model, and paired episodes and vary how strongly task state reaches the agent: a raw transcript, an exact checklist, per-turn directives from a state machine compiled from the brief and advanced only by execution receipts, or an enforcement gate on that machine that refuses state-violating actions; every episode is scored by exact payload matching against dynamic ground truth. Across three models, two reasoning regimes, and two domains, four findings hold without per-turn reasoning: displaying accurate state is unreliable, an unverified ledger the agent writes itself beats an accurate checklist it is shown, directives help in proportion to the model's obedience, and enforcement needs no obedience but is bounded by the correctness of its state and by the matcher that maps requests to steps; per-turn reasoning at a 235B agent compresses these separations without repairing the text rungs. The same gate, compiled from τ squared-bench's airline policy, raises a 235B agent's pass 1 from 0.39 to 0.54 and changes nothing for a 35B agent that rarely violates the policy; on PM-Bench, where acting turns on recognizing a cue rather than on state, showing the record is the best rung--matching or beating both gates and reversing the ledger-over-checklist finding--and enforcing the matcher's judgement drops a 35B agent below its raw transcript. Enforcement pays when failures are state-decidable and frequent, and hurts when the gate's judgement is wrong.
Sources
- AriGraph: Learning Knowledge Graph World Models with Episodic Memory for LLM Agents
- Constitutional AI: Harmlessness from AI Feedback
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Push Your Agent: Measuring and Enforcing Quantitative Goal Persistence in Long-Horizon LLM Agents
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- SABER: Small Actions, Big Errors -- Safeguarding Mutating Steps in LLM Agents
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- AgentLTL: A Trace-Verification Framework for Measuring, Enforcing, and Training Procedural Compliance in Tool-Using LLM Agents
- Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks
- Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following
- Anticipate and Learn: Unleashing Idle-Time Compute in Proactive Agents
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Deontic Policies for Runtime Governance of Agentic AI Systems
- A Human-Inspired Reading Agent with Gist Memory of Very Long Contexts
- User as Code: Executable Memory for Personalized Agents
- MemOS: A Memory OS for AI System
- PM-Bench: Evaluating Prospective Memory in LLM Agents
- Lost in the Middle: How Language Models Use Long Contexts
- Did You Forget What I Asked? Prospective Memory Failures in Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection