Dependency-Aware Trajectory Refinement for Efficient Multi-Turn Agent Fine-Tuning
cs.CL
Submitted: 2026-09-16
Updated: 2026-09-16
Comments: AACL 2026 Findings
License: http://creativecommons.org/licenses/by/4.0/
The gist: Multi-turn agent trajectories often contain redundant rounds (failed tool calls, parallel sub-queries, verification-only steps) that inflate both training and inference cost.
Terminology
Abstract
Multi-turn agent trajectories often contain redundant rounds (failed tool calls, parallel sub-queries, verification-only steps) that inflate both training and inference cost. We propose viewing each trajectory as a round-level dependency DAG that exposes which rounds are globally load-bearing for the final answer, and fine-tune agents on trajectories refined through this DAG. Given an LLM-annotated DAG, these edits are deterministic and interpretable, with optional rephrasing. Models trained on these refined trajectories consistently outperform those trained on the original trajectories at lower inference cost. Specifically, across four multi-modal QA benchmarks, our refinements improve downstream accuracy by up to 1.7,pp over vanilla SFT (and 5.7,pp over an LLM-deletion baseline) while reducing per-sample inference messages by up to approximately 40% and inference tokens by up to approximately 48%, translating to substantial savings in compute and serving cost. Code is available.
Sources
- Qwen3-VL Technical Report
- Instruction Mining: Instruction Data Selection for Tuning Large Language Models
- SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language Models
- Mind2Web: Towards a Generalist Agent for the Web
- Seeking and Updating with Live Visual Knowledge
- WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent
- MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines
- What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning
- AgentInstruct: Toward Generative Teaching with Agentic Flows
- Humanity's Last Exam
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Toolformer: Language Models Can Teach Themselves to Use Tools
- OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
- Chain of Draft: Thinking Faster by Writing Less
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- STaR: Bootstrapping Reasoning With Reasoning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering