VideoWeaver: Evaluating and Evolving Skills for Agentic Long Video Generation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "VideoWeaver: Evaluating and Evolving Skills for Agentic Long Video Generation".
Tom: VideoWeaver introduces an agent harness and benchmark designed to evaluate and evolve skills for long video generation, addressing the gap in understanding how general-purpose agent frameworks can handle complex,
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So, let’s talk about the title itself, "VideoWeaver: Evaluating and Evolving Skills for Agentic Long Video Generation." It immediately tells you this isn't just about generating clips; it’s about a system that learns how to build its own way of doing things for long video tasks.
Jane: Exactly. The authors are tackling the challenge of taking one simple instruction and turning it into a lengthy video by letting the agent compose its own sequence of skills, instead of just following a pre-set recipe.
Lu: They introduce this concept where agents create their own workflows by stitching together foundation skills, which is a significant step away from fixed pipelines that we’ve seen before.
Meng: When they say "agentic long video generation," I think that points to the need for sophisticated planning, not just execution of basic commands.
Lalam: For me, it means we are moving toward agents that can handle really intricate creative briefs without needing a human to write down every single step beforehand.
The paper's summary: Tom: Now let’s get into the summary of what they actually did in "VideoWeaver." Essentially, they built a benchmark with sixteen different task categories and two hundred eighty-five specific cases that cover text, image, audio, video, and combinations.
Jane: That benchmark is split into training sets and test sets plus three out-of-distribution categories to make sure they’re testing both normal performance and how well the agents can handle completely new situations.
Lu: The core idea is that instead of just judging the final product, they propose an agent-as-judge mechanism that checks both the plan execution trace and the finished video, using things like metadata and intermediate files as evidence for scoring.
Meng: That’s interesting because it addresses a real problem: knowing *where* in a long generation process something went wrong is crucial for debugging.
Lalam: It really points toward a future where AI can self-diagnose its own creative process, which is a huge cultural shift if it works reliably.
The paper's improvements: Tom: What’s really exciting about the proposed skill evolution algorithm is how it refines the agent’s abilities iteratively. It goes through three stages: first inferring an initial composition skill, then using an optimizer agent to refine that based on feedback, and finally merging skills across different categories into one comprehensive creator skill.
Jane: So they aren't just giving the agent a set of tools; they are giving it a way to actively learn and improve *how* to use those tools for long video tasks.
Lu: The mathematical notation they use, like that S = arg max Sk X x in Dtest k phi x, Aexec(x, Sk), shows a very structured way to optimize those skills based on what the agent actually does.
Meng: I see the practical implication here is that this self-refining loop means the system gets better at planning and sequencing its actions over many attempts.
Lalam: That ability to evolve its own skill set based on external critique suggests a level of procedural intelligence we haven't fully realized yet in current agent architectures.
Conclusion: Tom: So, to wrap up this discussion on "VideoWeaver: Evaluating and Evolving Skills for Agentic Long Video Generation," the main points are that agents can build their own workflows instead of following fixed pipelines, and they use an agent-as-judge to fix failures at any step of a long process.
Jane: This means we get much better diagnostic capabilities for complex generation tasks, and the skill evolution loop is what pushes those systems toward higher quality outputs by learning from feedback.
Lu: The implication is that foundation skills alone aren't enough; you absolutely need those high-level composition skills to orchestrate them effectively across different media types.
Meng: Practically speaking, this suggests that future video agents won't just be sequence executors, but true procedural directors capable of self-correction during long runs.
Lalam: I think the ability for these systems to autonomously refine their creation strategy based on feedback is what really matters for the broader impact on creative AI culture.
Zhejiang University · ByteDance
cs.CV
Submitted: 2026-06-06
Updated: 2026-10-01
Code: https://github.com/JianhuiWei7/VideoWeaver
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 83/100
The gist: VideoWeaver introduces an agent harness and benchmark designed to evaluate and evolve skills for long video generation, addressing the gap in understanding how general-purpose agent frameworks can
Key concepts
- Foundation Skills
- These are basic, self-contained capabilities that a language model has for tasks like generating video clips, synthesizing images, or understanding media. They act as the primitive building blocks that agents use to perform actions in the generation process.
- Composition Skills
- These are high-level procedural policies that define how foundation skills should be orchestrated together to achieve a long-horizon task, such as creating a complete video. They represent the 'recipe' or workflow an agent designs by combining multiple foundation skills.
- Agent-as-Judge
- This is a novel approach where another agent inspects both the execution trace (the steps taken) and the final video output to score its performance. This method provides evidence, using metadata and intermediate files, to ground the scores in concrete details rather than just subjective judgment.
- Skill Evolution Algorithm
- This is a three-stage process where an optimizer agent iteratively refines composition skills. It starts by inferring an initial skill, then uses feedback from the evaluation agent to optimize it, and finally merges category-level skills into a single creator skill for better performance.
Terminology
Summary
VideoWeaver introduces an agent harness and benchmark designed to evaluate and evolve skills for long video generation, addressing the gap in understanding how general-purpose agent frameworks can handle complex, long-horizon multimodal tasks by enabling agents to compose their own workflows rather than following predefined pipelines.
The gist
VideoWeaver is an agent harness and benchmark that evaluates and evolves skills for long video generation, where an agent turns a single instruction into a long video by composing foundation skills into its own workflow rather than following a predefined pipeline.
Benchmark Construction and Evaluation Methodology
The benchmark consists of 16 task categories and 285 cases,
spanning text, image, audio, video, and their combinations. The dataset is split into in-distribution train, test sets, and three out-of-distribution categories to assess both in-distribution performance and generalization. Evaluating these agents is difficult because errors can occur at any stage of the generation process,
necessitating a novel approach. To address this process-level failure diagnosis, the paper proposes an agent-as-judge that inspects both the execution trace and the final video,
grounding its scores in evidence such as metadata and intermediate files.
Skill Evolution Algorithm
The core innovation lies in refining agent capabilities through feedback. The proposed skill evolution algorithm is designed to progressively refine and merge skills. This process involves three stages:
-
Task Inference: The execution agent generates an initial category-level composition skill, denoted as
S0,k.
-
Skill Evolution: The optimizer agent iteratively refines the composition skill and creator skill using feedback from the evaluation agent to obtain optimized skills, such as
S⋆ = arg max Sk X x∈Dtest k ϕ x, Aexec(x, Sk).
-
Creator Skill Merging: The merge agent integrates experience across different categories into a single creator skill by progressively disclosing and merging category-level creator skills.
Metrics and Results
The evaluation utilizes two main sets of metrics: process metrics and output metrics. Process metrics check execution checkpoints in the trace, such as Planning PL
and Successful tool calls / total tool calls EE.
Output metrics assess quality, emphasizing cross-clip consistency,
which includes Visual Consistency VC
(characters and scenes remain consistent across clips) and Audio Consistency AC
(speaker identity and background music style remain consistent across clips). Experiments show that an explicit composition skill improves the generation process over using foundation skills alone,
while the skill evolution further improves output quality. Furthermore, a human study confirms that the proposed agent-as-judge aligns well with human judgments, especially on process metrics.
Foundation Skills and Composition Skills
The harness environment equips a backbone language model with a set of foundation skills, self-contained, independently invocable capabilities for video, image, and audio generation, understanding, and media processing.
These are formalized into foundation skills
(e.g., generating video clips or synthesizing images) that form the primitive action space. Beyond these isolated capabilities lies the concept of composition skills,
which are defined as high-level procedural policies that specify how foundation skills are orchestrated to accomplish such long-horizon generation tasks.
The system also defines a skill-creator
as a skill that constructs composition skills from available foundation skills and task cases.
Comparative Analysis
Experiments across multiple agent frameworks and foundation models reveal performance variations. Performance varies notably across harness and model choices,
with mature industrial settings achieving the strongest overall results on both process and output metrics. The comparison between No Composition Skill
and Composition Skill
demonstrates that foundation skills alone are insufficient, as the latter shows improvements in critical steps like Input Processing IP and Planning PL.
The final results indicate that Evolution with Feedback achieves a higher output average than Self-Evolution,
suggesting that external feedback provides valuable output metrics feedbacks harder for the agent to self-evaluate from its own execution.
Human Alignment and Limitations
The agent-as-judge is validated against human annotations, showing alignment scores are higher for process metrics and FR mainly because they are easier to evaluate than open-ended output quality.
Limitations include the current dataset size due to computational costs, the foundation skills being primarily from the ByteDance ecosystem, and future work focusing on more principled optimization methods
and stronger skill verification.
Potential risks involve intellectual property concerns regarding copyrighted assets and substantial computational cost.
Foundation Composition Skills
The set of foundation skills includes capabilities like:
(List extracted from Table 9)
(Examples include: AI Avatar Video Generate, One-Sentence Anime Video, Object Evolution Video, Cinematic Video, Audio-Driven Story, Long Video Edit.)
Foundation Composition Skills (Detailed)
The provided foundation skills expose basic video, image, audio, understanding, metadata,
and mediaprocessing capabilities.
For generation skills (e.g.
Improvements for AI systems
As a fastidious and diligent researcher, I have analyzed the VideoWeaver: Evaluating and Evolving Skills for Agentic Long Video Generation
paper. The core innovation lies in moving from fixed pipelines to agent-composed workflows, supported by an evidence-grounded evaluation loop (agent-as-judge) and a skill evolution algorithm.
Based on this research, here are specific improvements that can be made to AI systems, categorized by the capability they enable:
The improved AI system will operate as a highly autonomous, long-horizon creative director capable of generating complex cinematic narratives and personalized media from minimal high-level instructions. It moves beyond simple tool use to perform complex procedural planning and self-correction across multiple generative steps.
Here are the specific improvements:
-
A system that can turn a single high-level instruction (e.g.,
Create a 5-minute cinematic vlog about my trip to Tokyo
) into a complete, multi-shot video by dynamically composing necessary foundation skills into its own custom workflow, rather than following a pre-set sequence. -
The ability to perform long-horizon planning by decomposing the instruction into intermediate objectives, selecting and sequencing the appropriate foundation skills (video generation, image generation, audio synthesis), managing intermediate artifacts across these steps (e.g., character consistency data), and ensuring narrative, visual, and audio consistency across all resulting clips.
-
A robust evaluation mechanism that doesn't just check the final output quality but actively inspects the entire execution trace—including planning steps and tool calls—to diagnose failures at any intermediate stage (e.g., a flawed plan or a failed tool call).
-
An
Agent-as-Judge
that uses extracted evidence from the execution trace (metadata, intermediate files) to ground its scoring on process metrics (likeExecution Error-Free Rate
andClip Merging
) and output metrics (likeVisual Consistency,
which specifically checks for character consistency across clips). -
A self-refining skill acquisition loop where the agent uses the feedback from the Agent-as-Judge to progressively refine its composition skills (how it sequences foundation skills) and creator skills (how it structures the overall generation process), leading to higher quality outputs and better generalization to unseen tasks.
-
The ability to adapt its orchestration strategy based on learned experience: explicit composition skills are shown to improve generation over using foundation skills alone, indicating the system learns when a specific procedural skill is superior for a given task type.
-
Superior performance in complex, multimodal tasks where coherence across different modalities (e.g., ensuring the audio track matches the visual scene transitions) is critical.
This improved AI system can specifically:
-
Generate long-form, coherent videos (minutes to hours long) that adhere to complex narrative constraints.
-
Handle intricate cross-modal dependencies (e.g., matching a specific character's appearance across multiple shots while ensuring the background music style remains consistent).
-
Self-correct its generation process in real-time by inspecting execution logs and intermediate video frames, fixing planning errors or tool call failures before the final output is compromised.
-
Generalize to entirely new, unseen long video generation tasks by transferring learned orchestration knowledge from training data to out-of-distribution scenarios.
Abstract
Agentic long video generation requires planning, tool orchestration, and cross-clip coordination over a long horizon. Most existing video agents either rely on static, human-crafted workflows, which require substantial manual effort and poorly adapt across tasks, or iteratively refine the output of the current task without persistently distilling execution experience into reusable skills for future tasks. We introduce VideoWeaver, an agent harness and benchmark that evaluates and evolves skills for long video generation. Given a single high-level instruction, an agent dynamically composes foundation skills into its own workflow rather than following a predefined pipeline. We construct a benchmark of 16 task categories and 285 cases, with references spanning text, image, audio, video, and their combinations. We further propose an evidence-grounded agent-as-judge that inspects both the execution trace and the final video to diagnose process and output failures. Based on this feedback, our evolution algorithm progressively refines category-level composition and creator skills, allowing recurring experience to guide dynamically constructed workflows for unseen cases. Experiments show that explicit composition skills improve the generation process over foundation skills alone, while skill evolution further improves output quality and generalizes to unseen cases. Incorporating judge feedback yields additional gains, especially on output metrics, and the agent-as-judge aligns well with human, particularly on process metrics. Code is available at https://github.com/JianhuiWei7/VideoWeaver.
Sources
- VideoPhy: Evaluating Physical Commonsense for Video Generation
- AgentRx: Diagnosing AI Agent Failures from Execution Trajectories
- CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification
- Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
- LoL: Longer than Longer, Scaling Video Generation to Hour
- TRAIL: Trace Reasoning and Agentic Issue Localization
- Video-Bench: Human-Aligned Video Generation Benchmark
- PersonaVlog: Personalized Multimodal Vlog Generation with Multi-Agent Collaboration and Iterative Self-Correction
- EvoSkill: Automated Skill Discovery for Multi-Agent Systems
- VBench: Comprehensive Benchmark Suite for Video Generative Models
- XSkill: Continual Learning from Experience and Skills in Multimodal Agents
- LoViC: Efficient Long Video Generation with Context Compression
- SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?
- SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents
- Stable Video Infinity: Infinite-Length Video Generation with Error Recycling
- Let's Verify Step by Step
- VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning
- Memento-Skills: Let Agents Design Agents
- MetaClaw: Just Talk -- An Agent That Meta-Learns and Evolves in the Wild
- Evolving Medical Imaging Agents via Experience-driven Self-skill Discovery
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models