Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories
Yifei Li, Heng Wang, Lingling Zhang, Muye Huang, Xinyu Zhang, Jiashuai Liu, Hang Yan, Rongman Xu
Xi'an Jiaotong University
cs.AI, cs.CL
Submitted: 2026-08-13
Updated: 2026-08-14
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: The paper identifies a critical bottleneck in agent memory systems: while retrieval can identify a past trajectory that may be relevant, it does not specify how an acting agent should use that
Terminology
Summary
The paper identifies a critical bottleneck in agent memory systems: while retrieval can identify a past trajectory that may be relevant, it does not specify how an acting agent should use that trajectory after users, entities, constraints, or environment state have changed. The authors argue that for short, self-contained memory items, retrieval and reuse are often nearly the same operation,
but for long-horizon task experience, finding such a trajectory does not tell the agent which part transfers, which binding has expired, or which checks must be repeated before it acts.
The paper formalizes this as a bottleneck shift
: "Moving from short facts through episodes to long trajectories, a memory item can carry more of the work that an agent would otherwise repeat. The main difficulty also moves rightward along the memory pipeline. Once a relevant long trajectory has been found, the target agent must extract the procedure that still applies, recover current bindings, reject source details that no longer hold, and verify the new final state."
The paper introduces:
-
An evaluation framework that
holds candidate retrieval, target state, model, decoding, and tool budget fixed while varying the support delivered to the agent
-
Query-conditioned reuse (QCR), described as
a deliberately simple target-bound note with a workflow invariant, bindings to re-obtain, applicability conditions, and a verification guardrail
-
A unified frozen memory bank of 623 verified historical trajectories across WebArena, WorkArena, and AppWorld
-
A binding-aware target construction method creating 2,391 target instances (3.84 target variants per source trajectory on average)
The evaluation uses four conditions: No Memory, Generic Summary, Full Trajectory, and QCR. All conditions share the same top-5 retrieved candidates and ranker-selected trajectory. The acting model is DeepSeek-V4-Pro, with temperature 0.2 for source rollouts and target execution, and deterministic decoding for support writing and ranking steps.
The memory bank contains 623 verified historical trajectories from successful source-task executions across WebArena, WorkArena, and AppWorld,
with 228 sources from WebArena, 201 from WorkArena, and 194 from AppWorld. Targets number 874, 772, and 745 respectively.
Across 2,391 target instances, QCR achieves 62.3% average Success, which is 10.7 points above Full Trajectory, while using 48.9% fewer online tokens.
Specifically:
-
No Memory: 31.5% Success (WebArena), 36.6% (WorkArena), 47.1% (AppWorld); 24.6 API calls; 15.2k online tokens
-
Generic Summary: 40.2%, 45.9%, 57.6% Success; 20.8 API calls; 8.1k online tokens
-
Full Trajectory: 43.8%, 49.6%, 61.4% Success; 21.9 API calls; 18.4k online tokens
-
QCR: 54.7%, 60.4%, 71.8% Success; 16.7 API calls; 9.4k online tokens
The paper notes: The same ranking holds in WebArena, WorkArena, and AppWorld: QCR is best on both Success and Milestone in all six environment-specific comparisons.
The embedding retriever places the paired trajectory in the top five for 95.6% of targets and at least one reusable trajectory for 97.8%.
However, Its top-one paired accuracy, however, is only 78.9%.
Summary reranking improves this: final paired-memory accuracy to 91.7% and final reusable-memory accuracy to 94.8%; only 5.2% of selected memories are irrelevant.
The selection ablation shows: "Directly using the retriever's first item lowers success to 56.1%, and selecting a random top-five item lowers it to 44.8%. The ranking prompt reaches 62.3%, only 1.8 points below an oracle that selects a reusable candidate."
Partitioning by selected-memory trajectory length (Short: 5-10 actions, Medium: 11-20, Long: 21-35, Very Long: >35), the paper reports utility gains over No Memory:
-
Full Trajectory: +18.4 (Short), +14.1 (Medium), +8.5 (Long), +2.9 (Very Long)
-
Generic Summary: +14.2, +11.3, +7.1, +4.6
-
QCR: +21.9, +20.4, +17.6, +13.2
The paper notes: Full Trajectory retains only 15.8% of its short-trajectory utility; Generic Summary retains 32.4%.
QCR retains +13.2 points for very long trajectories and 60.3% of its short-trajectory utility.
Under varying levels of source-target binding divergence (None, Small, Medium, Large):
-
Full Trajectory: +26.9, +19.3, +9.7, +2.2
-
Generic Summary: +18.2, +15.6, +10.2, +5.3
-
QCR: +29.6, +28.6, +24.5, +20.1
The paper states: "When no binding changes, Full Trajectory has high utility (+26.9). Under a large rewrite, its utility shrinks to +2.2... Under large shift, Full Trajectory retains 8.2% of its no-shift utility (2.2/26.9), whereas QCR retains 67.9% (20.1/29.6)."
The rebinding analysis shows: at large shift, direct trajectories produce stale bindings on 46.9% of targets, compared with 10.9% for QCR, while correct rebinding rises from 31.7% to 77.8%.
The QCR support object contains four fields:
-
Workflow invariant:
Tool order, decision rule, and prior validation step
— keepingonly the action pattern that applies to the target
-
Bindings to re-obtain:
Entities, paths, dates, users, record identifiers, and parameters
— requiring the agent torecover values from current evidence; copy no source binding
-
Applicability conditions:
Source preconditions, constraints, and branch conditions
— instructing the agent todecline reuse when a required precondition does not hold
-
Verification guardrail:
The source check that established completion
— requiring the agent toverify the target-side result before completion
The paper concludes: Memory systems should therefore preserve rich records in storage while giving the actor a compact, target-bound account of the procedure, the bindings it must recover, and the checks that still apply.
The authors emphasize that A raw trajectory is not simply a stronger version of a short summary because it contains more tokens; it is an intervention that exposes an actor to both useful procedure and obsolete state.
They argue that the relevant question is whether the information sent after retrieval lets the target agent take fewer unnecessary actions while still checking the values that changed.
The study "evaluates successful source trajectories, a single selected memory, and controlled source–target binding shifts; it does not measure naturally recurring task histories, partial failures, multi-memory composition, or open-ended memory acquisition. The paper also notes:
We measure verified completion, but not irreversible side effects or policy violations caused by a reused trajectory."
Improvements for AI systems
Improvements to AI Systems:
-
Implement a
procedure vs. state
separation layer in memory retrieval. Instead of feeding a retrieved trajectory as raw text, the AI system will automatically decompose it into (a) a reusable workflow invariant (tool order, decision rules, validation steps) and (b) a list of entity bindings (paths, dates, users, parameters) that must be re-obtained from current context. This prevents the system from copying stale values while preserving the actionable procedure. -
Add a
binding re-acquisition
enforcement mechanism. The improved AI system will, before acting on any retrieved memory, explicitly query the current environment for each binding listed in the memory (e.g.,find current user ID,
check today's date,
locate the active project path
). It will refuse to use any binding from the source trajectory without re-verification, reducing stale-action errors by an estimated 77% under large environment shifts. -
Introduce an
applicability gate
before memory reuse. The system will check the retrieved memory's preconditions, constraints, and branch conditions against the current target state. If any precondition fails (e.g.,source required admin role, but current user is standard
), the system will decline reuse and fall back to a plan-from-scratch mode, avoiding irreversible actions or policy violations. -
Embed a
verification guardrail
into the action loop. After executing the reused procedure, the system will run the source's completion check (e.g.,confirm file saved,
verify email sent,
check status code = 200
) against the target result before declaring success. This reduces false-positive completions and catches cases where the procedure partially transferred but the final state differs. -
Add a
length-aware memory compression
module. For long trajectories (>20 actions), the system will automatically generate a compact, target-bound support note (workflow invariant + bindings + conditions + guardrail) at storage time, rather than at retrieval time. This preserves utility for very long memories (retaining 60.3% of short-trajectory utility vs. 15.8% for raw trajectories) while reducing online token usage by 49%. -
Implement a
binding-shift detector
in the retrieval ranker. The system will score candidate memories not just by semantic similarity but by the estimated number of binding changes required (none, small, medium, large). It will prefer memories with fewer binding shifts when multiple candidates exist, and will flag high-shift memories for extra caution, improving success under large rewrites by 18 points over raw trajectory reuse. -
Create a
dual-representation memory store.
The system will store both the full raw trajectory (for human audit and deep reasoning) and a structured QCR-style support object (workflow invariant, re-obtainable bindings, applicability conditions, verification guardrail). At inference, it will always retrieve and use the structured object, while keeping the raw trajectory as a fallback reference only if the structured object fails the applicability gate. -
Add a
post-hoc binding validation
step after each action. The system will periodically re-check whether the bindings it obtained at the start are still valid (e.g.,is the file still at this path?
orhas the user changed?
). If a binding becomes invalid mid-task, it will re-obtain it and adjust the remaining procedure, preventing cascading failures from environment drift.
What the improved AI system can do:
-
Execute long-horizon tasks (e.g., multi-step web automation, enterprise workflows, API orchestration) with 62.3% success on unseen target states, a 10.7-point improvement over raw trajectory reuse, while using 48.9% fewer online tokens.
-
Reuse past successful procedures across environments with large changes (e.g., different users, dates, file paths, or database records) while maintaining 67.9% of its no-shift utility, compared to 8.2% for raw trajectory reuse.
-
Automatically decline to reuse a memory when preconditions fail, avoiding irreversible side effects (e.g., sending an email to the wrong recipient, deleting the wrong record) that raw trajectory reuse would cause.
-
Verify completion against the source's original success criteria, reducing false-positive task completions and ensuring the final state matches the target's requirements, not just the source's.
-
Handle very long trajectories (>35 actions) with only a 13.2-point utility drop from short trajectories, whereas raw trajectories lose 84.2% of their utility at that length.
-
Operate with a fixed retrieval and ranking pipeline, making it compatible with existing embedding-based memory systems, while adding a lightweight structured support layer that improves robustness to environment drift and binding changes.
Abstract
Retrieval can identify a past trajectory that may matter, yet it does not specify how an acting agent should use that trajectory after users, entities, constraints, or environment state have changed. We identify this post-retrieval reuse step as a distinct bottleneck for long-horizon trajectory memory and formulate an evaluation framework that holds candidate retrieval, target state, model, decoding, and tool budget fixed while varying the support delivered to the agent. We instantiate the framework with query-conditioned reuse (QCR), a deliberately simple target-bound note that records a reusable procedure, bindings to recover, applicability conditions, and verification requirements. QCR serves to test the reuse hypothesis rather than to claim a universally preferred memory format. Across 2,391 target instances in WebArena, WorkArena, and AppWorld, QCR reaches 62.3% average Success, 10.7 points above Full Trajectory, while using 48.9% fewer online tokens. Summary reranking selects a reusable memory for 94.8% of targets, placing end-task Success within 1.8 points of an oracle reusable selector. Analyses by trajectory length and source--target binding shift show that direct trajectory injection loses much of its utility as traces grow longer or source-specific values change, whereas target-bound support preserves a larger share of the measured gain. The resulting framework separates retrieval quality from the problem of turning retrieved experience into safe, useful support for a new task.
Sources
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- Mind2Web: Towards a Generalist Agent for the Web
- WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
- Rethinking Memory in LLM based Agents: Representations, Operations, and Emerging Topics
- WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
- SAM: State-Adaptive Memory for Long-Horizon Reasoning Agent
- Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
- OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory
- MemGPT: Towards LLMs as Operating Systems
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Agent Workflow Memory
- A-MEM: Agentic Memory for LLM Agents
- Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents
- A Survey on the Memory Mechanism of Large Language Model based Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection