FRAMEWORKERS: A Dynamic Multi-Agent Framework for AI-Generated Video Production
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "FRAMEWORKERS: A Dynamic Multi-Agent Framework for AI-Generated Video Production".
Jane: The paper was written by Zhendong Li, Lei Sun, Letian Shi, Deheng Zhang, Ruibo Ming et al. from INSAIT Sofia University “St. Kliment Ohridski” and Snap Inc..
Tom: Stay tuned as we take you through the paper and discuss its implications.
Summary of the Mechanism: Tom: We just covered the high-level idea of "FRAMEWORKERS," and now the authors explain exactly how this entire closed loop functions from start to finish.
Jane: It’s a beautiful feedback system where every part of the process is constantly feeding information back into itself.
Meng: The whole thing operates through a Director, which decides what to do, and an Assistant, which executes the tasks and manage all the assets in a shared Workspace.
Lu: This workflow ensures that we aren't just running agents in isolation; they are all operating within a central repository of context and history.
Lalam: That persistent record is key to Lalam because it means that every choice made by the system is recorded, allowing us to see how the creative process evolved.
Tom: It’s a structured way of keeping track of everything, from the initial prompt to every single generated clip.
Jane: The Assistant takes those decisions and turns them into concrete actions by retrieving materials from that Workspace and passing them to the sub-agents.
Meng: From a practical standpoint, this is how we ensure that if an input brief requires a specific style image, the system knows exactly where to pull it from before sending it to the video generation model.
Lu: It allows us to manage the whole lifecycle of an asset—from its initial concept as text to its final form in a generated clip.
Lalam: This ensures that we are building a coherent narrative, not just a random collection of beautiful visuals.
Tom: We're seeing this entire structure, which provides the necessary foundation for the next step: looking at what makes this framework truly better than previous methods.
Improvements Over Existing Models: Tom: We’ve seen how FRAMEWORKERS is put together, and now let’s focus on why it's so much better than the models we used before.
Jane: The authors argue that older multi-agent systems are too rigid because they assume a fixed order for every task.
Meng: They point out that if those initial assumptions fail, the entire workflow breaks down, which is unacceptable in real production complexity.
Lu: The key improvement here is that you can modify the planned sequence without having to delete the history of everything you’ve already done.
Lalam: That ability to revise the plan dynamically allows us to adapt instantly when a story or an image doesn' change, mirroring how human creativity actually works.
Tom: It’s about having the intelligence to adjust mid-course, not just rigidly following a predetermined script.
Jane: So, if we needed a specific scene revised because of that dynamic planning ability, the framework could update its remaining tasks immediately incorporate that new information.
Meng: This adaptability is paired with modularity; since it uses semantic descriptions instead of hardcoded IDs, it’ can handle unexpected inputs.
Lu: That means a developer could introduce a brand new creative tool, like an advanced cinematic lighting effect, and the Director could figure out how to use it without rewriting the entire system.
Lalam: This is what makes the difference between a brittle proof-of-concept and an actual usable tool for creative teams.
Tom: We’re looking at a framework that doesn't just being powerful, but also designed to grow and adapt over time.
Jane: And that sets the stage perfectly for seeing how this theory holds up when we look at the hard data in their experiments.
Rigorous Testing and Results: Tom: The researchers put FRAMEWORKERS through rigorous testing using two demanding datasets, DCE (Easy) and DCH (Hard), to test its core planning capability.
Jane: We are looking specifically at how well the Director can actually plan a complex, multi-step sequence of actions.
Meng: They use metrics like Chain accuracy and Edit distance to show us exactly how close the predicted plan is to the ideal blueprint.
Lu: The results in Table one are impressive; the framework achieved ninety-one percent accuracy on the easier set and over eighty percent on the harder, more complex set.
Lalam: That level of consistency suggests that this system is not just good for simple tasks, but capable of handling deep narrative complexity.
Tom: Even more striking is its ability to handle failure; it recovers from runtime problems at a rate exceeding ninety-six percent.
Jane: It’s not perfect, but it’s incredibly robust when the parts break or something goes wrong in the middle of execution.
Meng: That high recovery rate is vital because any complex production pipeline will encounter inevitable technical hiccups during real-world operation.
Lu: This confirms that the dynamic nature of this planning is not just a theoretical neat trick but a practical, reliable mechanism for operational success.
Conclusion and Future Outlook: Tom: We’ve seen how FRAMEWORKERS handles complex video production tasks with such incredible reliability and adaptability across all those demanding tests.
Jane: It really feels like we are witnessing the establishment of a completely new standard for how automated content creation should function.
Lu: I think the real cultural shift here is that this enables us to create narratives that are not only technically flawless but also incredibly diverse in their execution.
Meng: From a practical standpoint, I’m excited about the ability scaling this framework into a complex industrial pipeline for production is a major step forward.
Lalam: It’s about taking vague creative concepts and turning them into tangible, consistent reality that gives artists unprecedented freedom over the final video.
Tom: That consistency is what makes it so powerful, especially when you consider the wide range of chaotic inputs users might throw at it.
Jane: We're really looking at a system that handles all the messy parts of making a film—the scripting, the asset management, and the final assembly—without ever breaking down.
Lu: It’s clear this moves beyond just being an AI tool; it provides a foundational orchestration engine for any kind of creative work.
Meng: The modularity means that this system can handle any type of input, whether it's a simple text prompt or a detailed reference image, without needing redesign.
Lalam: It ensures that the future-proof nature of this allows us to build stories with both deep consistency and massive variety in the final video output.
Tom: We’re looking at “FRAMEWORKERS: A Dynamic Multi-Agent Framework for AI-Generated Video Production” as a tool that is not only powerful but also incredibly robust.
Jane: It sets a high bar, showing us what’s possible when we move beyond rigid pipelines and embrace dynamic, smart workflows.
Lu: This is the kind of foundational thinking I hope sees more widespread adoption in all the AI systems we see today.
Meng: I just hope the engineering community uses this as a blueprint for of how to build complex reliable systems.
Lalam: It’s an exciting moment for automated storytelling and cultural expression, and it feels like a very important step forward for us.
Zhendong Li, Lei Sun, Letian Shi, Deheng Zhang, Ruibo Ming, Jian Wang, Danda Paudel, Luc Van Gool, Jinjin Gu, Mengshun Hu1, Dannong Xu1, Zhendong Li1, Lei Sun1, Letian Shi1, Deheng Zhang1, Ruibo Ming1, Jian Wang2, Danda Paudel1, Luc Van Gool1
INSAIT Sofia University “St. Kliment Ohridski” · Snap Inc.
cs.AI
Submitted: 2026-08-30
Updated: 2026-08-30
Project page: https://frameworkers.insait.ai
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 78/100
The gist: FRAMEWORKERS is a "task-centric and workspace-grounded multi-agent framework" designed to automate the complex, interdependent steps of AI-generated video production.
Key concepts
- FRAMEWORKERS
- This is a dynamic multi-agent framework for AI video production. It functions through a Director, which makes decisions, and an Assistant, which executes tasks by managing assets within a shared Workspace. This ensures that every choice and action is recorded in a central repository of context and history.
- Dynamic Planning
- Unlike rigid older systems, this framework allows users to modify the planned sequence without deleting the history of previous actions. This adaptability lets the system instantly adjust when a story or image changes, mirroring how human creativity works by allowing for mid-course corrections.
- Modularity
- The framework uses semantic descriptions instead of hardcoded IDs, making it highly adaptable to unexpected inputs. This allows developers to introduce new creative tools, such as advanced cinematic lighting effects, and enables the Director to figure out how to use them without needing a complete system rewrite.
Terminology
Summary
FRAMEWORKERS is a task-centric and workspace-grounded multi-agent framework
designed to automate the complex, interdependent steps of AI-generated video production. While modern generators excel at synthesizing individual clips, complete production requires long-horizon orchestration
of scripting, storyboarding, and editing. This framework matters because it replaces rigid pipelines
with a dynamic system capable of handling diverse user requirements and recovering from runtime failures.
The Multi-Agent Architecture
The framework organizes execution capabilities as a catalog of independent sub-agents
rather than a hand-crafted workflow. This modular design allows for plug-and-play capability expansion,
where new creative functions can be introduced by registering new sub-agents without redesigning the orchestration workflow. Each sub-agent is exposed through a descriptor specifying its:
-
Callable identifier and semantic role
-
Required input artifacts and produced output artifacts
-
Intended trigger conditions
Director and Assistant Roles
The system achieves separation between strategic planning and concrete execution
through two complementary modules. The Director acts as the high-level orchestration module, maintaining a Dynamic Task Stack
that represents the remaining work. It manages the workflow through two primitive operations: ADD(task, position) and DELETE(task id). This allows the Director to perform initial planning, task refinement, failure recovery, and adaptation to new user instructions.
The Assistant serves as the execution layer, grounding tasks in a shared Workspace
that contains four persistent components:
-
file system: Stores textual and structured records such as briefs and scripts. -
generated assets: Stores multimedia artifacts like images, videos, and audio. -
global memory: Maintains high-level summaries of completed steps and decisions. -
logs: Preserves detailed execution traces for debugging and recovery.
Training and Orchestration Reliability
To improve orchestration reliability,
the Director is fine-tuned using a two-stage process. First, supervised fine-tuning (SFT) teaches descriptor-conditioned task routing
by supervising structured rationales and executable plans. This is followed by Group Relative Policy Optimization (GRPO) to optimize for valid output format, complete sub-agent coverage, and consistent ordering.
This specialized training enables the Director to reason over long-range dependencies, artifact availability, and the functional constraints of heterogeneous sub-agents.
By optimizing these properties, the framework can recover reliably from runtime failures
and generalize to unseen sub-agents without retraining.
Performance and Capabilities
Experiments demonstrate that FRAMEWORKERS outperforms strong LLM planners in routing accuracy
and achieves higher end-to-end video quality and broader task coverage
than fixed pipelines. The framework's modularity provides several key advantages:
-
It supports a
broad range of AIGC video production tasks
from diverse specifications. -
It maintains superior
entity consistency
across character, prop, and scene elements. -
It achieves a
96.7% overall recovery rate
when encountering plan-level or execution-level perturbations.
Improvements for AI systems
1. Dynamic Task-Stack Orchestration
-
Improvement: Replace rigid, linear pipelines with a non-linear, editable Dynamic Task Stack managed by a central Director.
-
System Capability: The AI can perform mid-execution replanning, allowing it to insert recovery tasks (e.g.,
regenerate character design
) or delete obsolete tasks in response to changing user instructions or intermediate failures, rather than crashing or following a broken sequence.
2. Semantic Descriptor-Based Sub-Agent Integration
-
Improvement: Transition from hard-coded API calls to a descriptor-driven registration system where sub-agents are defined by their semantic roles, required inputs, and produced outputs.
-
System Capability: The AI achieves zero-shot extensibility, meaning it can autonomously integrate and utilize entirely new, previously unseen specialized tools (e.g., a new 3D-model generator or a specific style-transfer agent) simply by reading their functional descriptions, without requiring retraining of the central planner.
3. Workspace-Grounded Global Memory
-
Improvement: Implement a centralized, persistent Workspace (comprising a file system, generated asset repository, and global memory) to serve as the single source of truth for all agents.
-
System Capability: The AI can maintain strict long-horizon consistency, specifically preventing
identity drift
by ensuring that character facial attributes, clothing, prop geometry, and environmental layouts are explicitly retrieved and passed between disparate sub-agents (e.g., from aCharacter Design
agent to aVideo Generation
agent).
4. GRPO-Optimized Dependency Reasoning
-
Improvement: Fine-tune the orchestration model using Group Relative Policy Optimization (GRPO) specifically to reward sub-agent coverage and dependency-consistent ordering.
-
System Capability: The AI can reliably execute complex, multi-stage production chains that require strict logical prerequisites (e.g., ensuring a
Screenplay
is finalized and aStoryboard
is approved before anyVideo Generation
tasks are triggered), significantly reducing routing errors in long-horizon workflows.
5. Closed-Loop Evaluation and Replanning
-
Improvement: Integrate an automated
Evaluation-to-Replanning
loop where an Evaluator agent provides structured feedback on the quality and consistency of generated artifacts. -
System Capability: The AI can perform autonomous quality control; if an evaluator detects an error (such as a mismatch between a character's appearance in shot A vs. shot B), the system automatically interprets the failure and modifies the Task Stack to execute a corrective sub-task, ensuring the final output meets high-fidelity standards without human intervention.
Sources
- Movie Gen: A Cast of Media Foundation Models
- Automated Movie Generation via Multi-Agent CoT Planning
- VideoMemory: Toward Consistent Video Generation via Memory Integration
- Camera Artist: A Multi-Agent Framework for Cinematic Language Storytelling Video Generation
- Mora: Enabling Generalist Video Generation via A Multi-Agent Framework
- CineAGI: Character-Consistent Movie Creation through LLM-Orchestrated Multi-Modal Generation and Cross-Scene Integration
- StoryAgent: Customized Storytelling Video Generation via Multi-Agent Collaboration
- Animate-A-Story: Storytelling with Retrieval-Augmented Video Generation
- AesopAgent: Agent-driven Evolutionary System on Story-to-Video Production
- UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist
- Position: Agentic Systems Constitute a Key Component of Next-Generation Intelligent Image Processing
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection