CoEvoWhen: Policy-Tool Coevolution for Ultra-Long Video Temporal Grounding
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "CoEvoWhen: Policy-Tool Coevolution for Ultra-Long Video Temporal Grounding".
Jane: Ultra-long video temporal grounding requires balancing long-range evidence search with fine-grained event understanding under a limited visual budget,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Let's talk about the title and authors of this paper, "CoEvoWhen: Policy-Tool Coevolution for Ultra-Long Video Temporal Grounding." It clearly tells us the core mechanism is about evolving both policies and tools together for ultra-long video grounding.
Jane: The authors are Yiduo Jia, Muzhi Zhu, Jinchuan Shi, Hao Zhong, Yuling Xi, Ke Liu, and Hao Chen from Zhejiang University and the State Key Lab of CAD and CG. They're the ones driving this policy-tool coevolution framework.
Lu: I find it compelling that they’re focusing on ultra-long video; that's where current methods really struggle with balancing searching far out while still getting fine details right.
Meng: It suggests a departure from just tuning a single model, as they are building this reusable skill externally using trajectories to refine the system.
Lalam: This approach points toward building skills that can be modularly plugged into larger systems, which is something we've been thinking about for improving cultural understanding of complex data.
The paper's summary: Tom: So, in terms of what this paper summarizes, CoEvoWhen introduces a framework where an external skill evolves by distilling experience from the VLM's reasoning paths. It starts with a minimal base skill and then uses an external skill updater to refine both the high-level policies and the executable media tools.
Jane: The summary emphasizes that for policy evolution, they distill transferable experience to improve how the agent plans its search and organizes its observations, while for tool evolution, they use coding capabilities to upgrade or create new media processing tools.
Lu: It’s interesting how they pair these two updates; the idea is that the policy determines what evidence is needed while the tools dictate exactly what kind of evidence can be presented back to the system during observation.
Meng: They explicitly show this interdependence, stating that tool-use policies and tool design are linked, which makes sense because you need good tools to execute a good plan for long video tasks.
Lalam: This structure suggests that the skill isn't just a set of instructions but a dynamic capability that learns how to balance searching globally with zooming in on specific details.
The paper's improvements: Tom: Now, looking at the improvements they propose, CoEvoWhen shows how this coevolution actually works by refining the policies and tools based on execution trajectories. They show that policy updates refine strategies for planning and observation orchestration, while tool updates adapt or expand media processing capabilities.
Jane: The core mechanism is formalized as updating the skill components: Policy (k+one) is refined by (k), and Tools T(k+one) are refined by T(k), which is how they manage that iterative improvement process.
Lu: The paper highlights a crucial insight about coordination, suggesting that comparing temporally distant candidates benefits from image-based observations for compact coverage, whereas judging action order needs local video-based observations to check continuity.
Meng: This points toward a coordinated observation strategy where the system knows exactly when to switch between global image search and detailed video verification based on what the query demands.
Lalam: The result of this coordination is a skill that supports image-based search, candidate refinement, and selective video verification, which seems like a very targeted way to use visual resources.
Conclusion: Tom: To wrap things up with the conclusion of "CoEvoWhen," it really boils down to how this policy–tool coevolution framework successfully evolves an external skill from a minimal base into something capable of ultra-long video temporal grounding.
Jane: The paper concludes that this method improves grounding accuracy while simultaneously reducing the visual token cost required for those complex tasks, showing a trade-off that is quite efficient.
Lu: The implication here is that by treating the skill as an evolving entity rather than just a static prompt, we can get better performance on challenging temporal reasoning without needing to update the underlying VLM parameters.
Meng: From a practical standpoint, this means we can deploy agents with much more sophisticated temporal understanding on longer video sequences without having to overhaul the massive model architecture itself.
Lalam: This reusable skill concept is powerful because it means we build expertise once and reuse it across different long-video tasks, which really makes scaling AI applications more feasible.
Yiduo Jia, Muzhi Zhu, Jinchuan Shi, Hao Zhong, Yuling Xi, Ke Liu, Hao Chen
Zhejiang University
cs.CV
Submitted: 2026-09-30
Updated: 2026-09-30
Comments: Project page: https://aim-uofa.github.io/CoEvoWhen/
Project page: https://aim-uofa.github.io/CoEvoWhen
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 90/100
The gist: Ultra-long video temporal grounding requires balancing long-range evidence search with fine-grained event understanding under a limited visual budget, yet existing agentic methods still rely largely
Key concepts
- Policy–Tool Coevolution
- This is the core mechanism where high-level policies (planning and orchestration) and executable media tools (processing capabilities) improve together. The system learns by distilling experience from how the policy guides tool use, refining both components simultaneously to create a reusable skill.
- Skill Representation
- The external skill is formally represented as two parts: high-level policies ($\Pi(k)$) that guide task planning and observation orchestration, and executable media tools ($T(k)$) that perform the actual media processing. These components are updated iteratively through experience to become more effective.
- Coordinated Observation Strategy
- The framework recognizes that tool design and policy are interdependent. It coordinates image-based search for long temporal ranges with video-based verification for local details, allowing the system to strategically choose between modalities based on whether it needs global evidence or fine-grained action order information.
- Inference at Test Time
- At runtime, the fixed evolved skill is applied to the frozen VLM. The policy guides tool invocation based on the query and accumulated evidence, while tools process that evidence. This allows the VLM to autonomously acquire necessary visual data in a structured way.
Terminology
Summary
Ultra-long video temporal grounding requires balancing long-range evidence search with fine-grained event understanding under a limited visual budget, yet existing agentic methods still rely largely on predefined policies and tool capabilities. The proposed framework jointly evolves high-level policies and executable media tools from the agentic reasoning trajectories of a VLM, forming a reusable skill without updating model parameters.
How it works
The core of CoEvoWhen is a policy–tool coevolution framework that iteratively improves an external skill, which consists of high-level policies and executable media tools. This process starts with a minimal base skill
providing only the basic grounding protocol, initial observation instructions, and primitive image and video observation operations.
The frozen Vision-Language Model (VLM) executes tasks using this evolving skill autonomously at inference time without needing a separate, stronger planning model.
Skill Representation and Coevolution Mechanism
At each evolution round, the external skill is represented as a set of high-level policies, denoted as Π(k),
and executable media tools, denoted as T(k).
The policy guides task planning and observation orchestration—specifying what visual evidence is needed—while the tools provide the executable media processing capabilities. The skill update mechanism involves distilling transferable experience from execution trajectories to refine these components:
-
For policy evolution,
∆Π(k) refines the strategies for task planning and observation orchestration with reusable experience distilled from the feedback.
-
For tool evolution,
∆T(k) adapts and expands the media processing capabilities available to the model by upgrading existing executable tools or creating new ones.
The formal update is expressed as: Π(k+1) = Π(k) ⊕ ∆Π(k), T(k+1) = T(k) ⊕ ∆T(k).
Coordinated Observation Strategy
A key insight of the framework is that tool-use policies and tool design are interdependent.
High-level policies determine what evidence is needed, when to acquire it, and how to organize the search,
while media tools dictate the form and granularity of the evidence the tools can present.
The system coordinates image-based search with video-based verification:
: "comparing temporally distant candidates calls for image-based observations that compactly cover long temporal ranges while preserving their temporal correspondence, whereas judging action order, state changes, and event continuity may require local video-based observations."
The resulting skill supports image-based search, candidate refinement, and selective video verification,
allowing the VLM to acquire evidence autonomously.
Evaluation and Results
Systematic evaluations across five benchmarks (VUE-LVTR held-out set, ExtremeWhenBench, CoMET-Bench) and three VLMs demonstrate that coevolution improves grounding accuracy while reducing visual token cost.
Key findings include:
-
Performance gains: On ExtremeWhenBench, evolving the base skill for Qwen3.5-27B
raises mIoU by 74.9%,
and on CoMET-Bench, it yields gains of0.0280 in mIoU and 10.93 points in Rejection-F1.
-
Efficiency improvements: Average visual tokens per query drop from 202.6k to 141.9k on the VUE-LVTR held-out set, with reductions of
30% fewer
tokens overall, demonstrating a reallocation of the visual budget between modalities. -
Generalizability: The evolved skill shows
cross-VLM generalizability,
improving ExtremeWhenBench mIoU by99.2%
and transferring effectively to general long-video QA without additional task-specific evolution.
Ablations and Insights
Ablation studies confirm the necessity of joint adaptation:
: freezing the policies during evolution leads to a larger accuracy loss than freezing the tools,
suggesting that stronger observation capabilities can yield substantial gains in ultra-long video temporal grounding when strategically organized into an efficient and accurate evidence search process.
Image–video coordination ablations show that the evolved skill outperforms single-modality skills, underscoring that image-based and videobased observations are suited to different contexts
and that strategically coordinating the two makes better use of their complementary strengths.
The final evolved skill successfully combines image-based global search with selective video verification.
Inference at Test Time
At inference, the fixed evolved skill is applied to the VLM. The process is formalized as: at = (jt, ηt) ∼ Mθ (· qi, Ht−1; Π∗, T∗), Ht = Ht−1 ∥ at, f ∗ jt(Vi; ηt).
The VLM autonomously invokes tools based on the query and accumulated evidence guided by the evolved policy.
Improvements for AI systems
Based on the provided research paper COEVOWHEN: POLICY-TOOL COEVOLUTION FOR ULTRA-LONG VIDEO TEMPORAL GROUNDING,
here are specific, actionable improvements for AI systems, categorized by capability:
) 1. Enhanced Temporal Grounding Accuracy in Ultra-Long Videos (Extreme Precision)
The improved system can achieve significantly higher temporal localization accuracy than current methods by dynamically adapting its observation strategy based on the evidence needs of the query.
Specific Improvements:
-
Implement a policy that explicitly coordinates image-based global search with fine-grained, selective video verification. This allows the system to perform
Global Image-Based Search
for distant candidates while reserving high visual cost video inspection only for local boundary refinement or action verification on promising regions (as demonstrated in Figure 6). -
Evolve media tools to support richer observation modalities. Tools should be upgraded to generate compact temporal summaries, motion-aware frames, and scene boundary indicators, enabling the system to distinguish between visually similar but temporally distinct events.
Improved AI System Capability:
The improved system can accurately localize target events in videos spanning hours (e.g., 76 minutes) by intelligently balancing long-range evidence search with fine-grained local detail inspection, leading to near-perfect temporal overlap (IoU of 1.00 demonstrated).
) 2. Significant Reduction in Visual Token Cost and Inference Latency
The system can achieve high accuracy at a substantially lower computational cost per query compared to current agentic methods that rely on monolithic, exhaustive visual scanning.
Specific Improvements:
-
Employ
efficient tool
policies that minimize redundant observation calls by leveraging pre-computed or summarized evidence (e.g., contact sheets) before initiating expensive video inspections. -
Evolve tools to be more compact and targeted, ensuring that image observations are used for broad candidate pruning while video observations are strictly reserved for verifying the specific temporal boundaries of the final answer, drastically reducing total visual token consumption (up to 51.27% reduction shown in Case Study).
Improved AI System Capability:
The system can perform complex ultra-long video grounding tasks with a 30% reduction in cumulative visual token cost per query, making high-accuracy understanding feasible for real-time or resource-constrained applications.
) 3. Autonomous, Self-Optimizing Agentic Reasoning (No External Planner Needed)
The system gains the ability to operate autonomously under a fixed, frozen VLM backbone without requiring a separate, stronger planning model for multi-step reasoning.
Specific Improvements:
-
Implement an
evolved skill
that integrates high-level policy guidance directly into the VLM's system prompt and registers evolved tools as callable functions. -
The skill updater (e.g., Codex) autonomously distills transferable task experience from execution trajectories to refine both the agent's planning strategy (policy) and its execution capabilities (tools).
Improved AI System Capability:
The agent can autonomously decide the next observation step—whether to search globally, refine a specific candidate using images, or verify dynamics using video—based on accumulated evidence and refined task experience, effectively acting as its own superior planner for long-video tasks.
) 4. Cross-Task Reusability and Generalization
The learned expertise is not task-specific but transferable across different types of long-video understanding problems (e.g., grounding vs. general QA).
Specific Improvements:
-
The coevolution process is designed to distill a
reusable skill
rather than just a solution for one query type. -
The framework demonstrates that the evolved grounding skill can be directly applied to general long-video QA tasks (like LVBench or LSDBench) without requiring any additional task-specific evolution, proving the robustness and generality of the learned experience.
Improved AI System Capability:
The system can rapidly adapt its learned temporal grounding expertise to new, related video understanding challenges with minimal or zero retraining on new data, showcasing a highly generalizable knowledge base for long-video comprehension.
) 5. Dynamic Adaptation to Query Complexity (Multi-Event Handling)
The system can effectively handle complex queries involving multiple events or target absences within the same video context.
Specific Improvements:
-
The policy is refined to handle
dynamic ambiguities
anduncertain continuity
by explicitly targeting action order, state changes, and event transitions during video inspection. -
Tools are adapted to provide evidence that helps discriminate between similar but distinct visual patterns (e.g., distinguishing a chute transfer from a skimmer action).
Improved AI System Capability:
The system can accurately count multiple events in a single video and correctly reject queries where the target event is absent, significantly boosting performance on benchmarks like CoMET-Bench.
Abstract
Ultra-long video temporal grounding requires balancing long-range evidence search with fine-grained event understanding under a limited visual budget, yet existing agentic methods still rely largely on predefined policies and tool capabilities. Motivated by this, we propose a novel policy-tool coevolution framework that jointly evolves high-level policies and executable media tools from the agentic reasoning trajectories of a VLM, forming a reusable skill without updating model parameters. During evolution, an external skill updater distills transferable task experience in long-video temporal grounding, accordingly refining the orchestration of long-range image-based and fine-grained video-based observations. Alongside these policy updates, the updater employs its coding capabilities to upgrade existing tools or create new ones, adapting the tools to long-video evidence acquisition. Equipped with the evolved skill, the VLM autonomously orchestrates tools under the guidance of the evolved policy, coordinating image and video observations for agentic inference without relying on a separate, stronger planning model. Extensive experiments spanning five benchmarks and three VLMs show that policy-tool coevolution consistently improves temporal grounding accuracy in ultra-long videos while reducing visual token cost at inference, and that the evolved skill yields substantial performance gains on general long-video QA without additional task-specific evolution, demonstrating the effectiveness and generalizability of our framework for long-video understanding.
Sources
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- XSkill: Continual Learning from Experience and Skills in Multimodal Agents
- EvoGround: Self-Evolving Video Agents for Video Temporal Grounding
- Natural-Language Temporal Grounding in Hour-Long Videos is a Search Problem: A Benchmark and Empirical Decomposition
- Vidi: Large Multimodal Models for Video Understanding and Editing
- Vidi2.5: Large Multimodal Models for Video Understanding and Creation
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Agent Workflow Memory
- SkillSmith: Co-Evolving Skills and Tools for Self-Improving Agent Systems
- Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence
- Conditional Multi-Event Temporal Grounding in Long-Form Video
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models