DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement
cs.AI
Submitted: 2026-09-18
Updated: 2026-09-26
Comments: 42 pages, including appendices
License: http://creativecommons.org/licenses/by/4.0/
The gist: Online agent deployments produce abundant execution traces, while task-specific verification and expert annotation are costly to scale.
Terminology
Abstract
Online agent deployments produce abundant execution traces, while task-specific verification and expert annotation are costly to scale. We study how to distill these traces into reusable feedback without post-hoc outcome labels, drawing on their evidence of local progress, recovery, and unfinished requirements. We introduce DENSE (Distilling Evidence from Nested Subtask Executions), which organizes this evidence into evidence-grounded nested shortcut trees. DENSE compresses redundant attempts, reconciles issues across levels using recovery evidence, and summarizes completed branches while expanding unresolved ones, linking reusable progress to remaining obligations. We introduce REFIT, a source-paired protocol comparing feedback from shared initial trajectories under post-hoc outcome blindness, with environments and model contexts reset for fresh attempts at the same tasks. On Terminal-Bench 2.1, DENSE achieves the highest strict pass rate among tested non-privileged feedback methods across four recipient models. Relative to initial executions, strict pass rate improves by 7.12-15.64 pp, with 19.0-43.6% fewer observed recipient tokens in reruns. GPT-5.5 ablations support combining nested subtask analysis with shortcut construction and issue reconciliation. These findings point toward agent self-refinement through evidence-grounded trajectory reuse with less reliance on external supervision.
Sources
- AgentRx: Diagnosing AI Agent Failures from Execution Trajectories
- The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
- Long-Horizon Agent Trajectory Attribution: A Unified Benchmark and Fine-Grained Annotation Framework
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- AgentProcessBench: Diagnosing Step-Level Process Quality in Tool-Using Agents
- On the Reliability of Computer Use Agents
- Falsification, Not Exposure: An Internally Preregistered Placebo-Controlled Decomposition of Self-Repair Feedback in Frozen Small Code Models
- DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems
- A Self-Evolving Framework for Efficient Terminal Agents via Observational Context Compression
- R$^3$L: Reflect-then-Retry Reinforcement Learning with Language-Guided Exploration, Pivotal Credit, and Positive Amplification
- Conversation for Non-verifiable Learning: Self-Evolving LLMs through Meta-Evaluation
- RealCritic: Towards Effectiveness-Driven Evaluation of Language Model Critiques
- Tongyi DeepResearch Technical Report
- Fail-Fast, Restart-Smart: Early Failure Prediction and Restart for SWE Agentic Tasks
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
- WebArena: A Realistic Web Environment for Building Autonomous Agents
- AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection