Scoped Verification for Reliable Long-Horizon Agentic Context Evolution under Distribution Shift
cs.AI, cs.CL
Submitted: 2026-07-10
Updated: 2026-09-21
Comments: 18 pages, 3 figs
Code: https://github.com/RedMind-Research/GRACE
License: http://creativecommons.org/licenses/by/4.0/
The gist: Deployed LLM agents rely on agentic context, the model-external textual control content assembled by an operational harness.
Terminology
Abstract
Deployed LLM agents rely on agentic context, the model-external textual control content assembled by an operational harness. In this work, the mutable component of that context is a persistent system-level instruction that is updated from operational experience while the model, tools, and harness remain fixed. Over long evolution horizons, flat-text maintenance makes verification increasingly difficult as accumulated instructions grow and interact. We propose Graph-Regularized Agentic Context Evolution (GRACE), which maintains the persistent instruction component as a typed semantic graph and validates proposed updates within the local typed neighborhoods of modified nodes. Accepted graph updates are reconstructed as incremental edits to the textual instruction checkpoint used at deployment. We evaluate GRACE within a fixed telecom agent harness derived from τ squared-bench under a controlled distribution-shift protocol. Across five independent replications, GRACE improves strict reliability, measured by pass cubed, from the Gemini 2.5 Flash zero-shot value of 0.091 to 0.673 plus or minus 0.136 at the final checkpoint. This exceeds a Gemini 3.1 Pro zero-shot reference of 0.242 on the same held-out set, while the flat-text HCE baseline finishes at 0.191 plus or minus 0.051. These results identify two requirements for reliable long-horizon context evolution, a structural substrate that makes verification local and a consolidation mechanism that keeps accumulated instruction content usable.
Sources
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- Meta-Harness: End-to-End Optimization of Model Harnesses
- The World Won't Stay Still: Programmable Evolution for Agent Benchmarks
- GEM: A Gym for Agentic LLMs
- SCOPE: Prompt Evolution for Enhancing Agent Effectiveness
- EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection