OmniWeaving: Towards Unified Video Generation with Free-form Composition and Reasoning

arXiv:2603.24458 · cs.CV · Submitted 2026-03-25 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "OmniWeaving: Towards Unified Video Generation with Free-form Composition and Reasoning".

Tom: OmniWeaving proposes an omni-level video generation model designed to achieve unified, free-form video creation by integrating powerful multimodal composition and reasoning capabilities.

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So Jane, we're looking at this paper called "OmniWeaving: Towards Unified Video Generation with Free-form Composition and Reasoning." The main idea here is that they've built a model that can handle video generation in a much more flexible way than what we see in current open-source models. They claim it achieves unified, free-form video creation by integrating strong multimodal composition and reasoning abilities.

Jane: That sounds really ambitious, Tom. So, if I understand correctly, the core thesis of OmniWeaving is to solve the problem where existing models are fragmented across different tasks—like just doing text-to-video or image editing separately—by creating a single framework that can handle everything together.

Lu: What's fascinating about their proposal is that they aren't just tacking on more features; they’re proposing a unified architecture built around an MLLM, an MMDiT, and a VAE to handle the visual comprehension and generation side simultaneously.

Meng: A unified architecture sounds neat in theory, Lu. But for us engineers, the real question is how they manage that complexity when you're trying to get something stable and actually runnable on hardware.

Lalam: From an engineering standpoint, I think what excites me most is the idea of this system acting like an intelligent agent to infer complex user intentions for video creation. That level of understanding could really improve how we design interactive AI experiences in the future.

Tom: Exactly, Lalam! It moves beyond just following explicit instructions; it suggests the model can actually reason about what someone *wants* to create, which is a big step forward. Jane, what's your take on why this unified approach is necessary right now?

Jane: Well, Tom, I think the necessity comes from the fragmentation we see in open-source models. They often struggle to seamlessly integrate diverse inputs like text and multiple images into one coherent video narrative without breaking down into separate steps. OmniWeaving aims to bridge that gap by temporally binding these different modalities together intelligently.

Lu: And they've achieved this by introducing specific enhancements to their model, like activating a thinking mode in the MLLM, which lets it generate intermediate reasoning steps to create a more precise prompt before conditioning the main generator. That’s really smart for boosting its composition power.

Meng: Reasoning steps sound powerful, Lu, but how much computational overhead does that thinking process add to the inference time? We need efficiency if we want this to be practical beyond just research papers.

Lalam: I think those intermediate reasoning steps are crucial because they allow the model to build a richer semantic understanding before it starts generating the video itself. That richer understanding could translate into much more nuanced and contextually appropriate outputs for users.

Paper summary: Tom: That's a great point, Lalam. It’s not just about speed; it’s about quality of intention, which leads us right into the training data they used to build this model—it sounds like they trained it on something massive.

Jane: They are training OmniWeaving on a massive-scale pretraining dataset that covers a wide spectrum of scenarios, including both real-world and synthetic domains, which gives it exposure to very diverse appearances and motions. This breadth is what allows it to handle those complex compositions the abstract paper mentions.

Lu: I think the structure of their training tasks really highlights how they tackled these challenges, focusing on foundational generation, multimodal composition with interleaved inputs, and reasoning-augmented tasks where it has to deduce temporal progressions from keyframes.

Meng: When you talk about deducing temporal progression from disparate keyframes, Lu, that’s where I get practical. Can the model actually handle the ambiguity of what happens *between* those frames in a way that doesn't just result in flickering or nonsensical motion?

Lalam: The paper suggests they tackle this by explicitly designing tasks like Event-Deductive Multi-Image-to-Video generation, which forces the model to create a reasoning trace detailing the temporal progression. This trace itself is what guides the video generation, giving it structure where there wasn't any explicit instruction.

Tom: So it’s not just guessing between frames; it’s using that inferred reasoning trace as a blueprint for motion. It sounds like they are moving toward a system that can handle truly free-form composition, which is what the title promises with OmniWeaving: Towards Unified Video Generation with Free-form Composition and Reasoning.

Jane: And the structure of this entire framework, integrating the MLLM's semantic understanding directly into the MMDiT’s generation process via hidden states deepstacking, really shows how they unified those two critical components. It makes sure that what the language understands directly informs how the visuals are put together.

Lu: That deepstacking technique is key because it injects multi-granular semantic guidance into the generative process, meaning different levels of semantic understanding from the MLLM can influence different parts of the video creation simultaneously. It’s a sophisticated way to condition generation.

Meng: I'm still thinking about how they are actually running this thing in a production environment. The paper focuses heavily on the architecture, but what are the practical implications for deploying something that requires such complex reasoning traces?

Lalam: I think the implication for culture is that if we can build systems that can reason about and synthesize such complex narrative structures from simple inputs, it opens up entirely new avenues for creative expression. It suggests a future where video creation isn't just about clicking buttons but about articulating abstract concepts in a way the AI understands deeply.

Paper summary: Tom: That’s a big picture idea, Lalam. The paper is clearly aiming for that level of capability by proposing this comprehensive system. Now, let's look at what they are actually trying to prove with their evaluation method.

Jane: Right, Tom, we need to talk about the IntelligentVBench benchmark they introduced because it seems like the authors recognized that existing benchmarks weren't sufficient for testing truly unified systems. They propose a "VLM-as-a-judge" paradigm for evaluating abstract reasoning and composition across four specific tasks.

Lu: That VLM-as-a-judge approach is significant because it moves the evaluation beyond simple metrics to assess the actual quality of the abstract reasoning and compositional skills, which is exactly what you need for a system that claims omni-capability.

Meng: From an engineering standpoint, having a novel benchmark like IntelligentVBench gives us something concrete to measure against when we try to build or compare our own unified systems. It forces us to define what 'omni-capable' actually means in terms of measurable performance on complex tasks.

Lalam: I think the VLM-as-a-judge is particularly interesting because it lets the AI judge its own output based on a high-level understanding, which could lead to systems that are self-correcting in their compositional choices over time. That kind of internal feedback loop could be very powerful for improving culture and creativity.

Tom: So, to wrap up this overview of "OmniWeaving: Towards Unified Video Generation with Free-form Composition and Reasoning," we've seen how they propose a unified architecture using the MLLM, MMDiT, and VAE to handle complex composition and reasoning.

Jane: And the authors are pushing for a more rigorous way to test this capability by introducing IntelligentVBench, which uses a VLM-as-a-judge paradigm to assess abstract reasoning across tasks like Implicit I2V and Compositional MI2V.

Lu: The overall claim is that by leveraging a massive dataset and this three-stage training strategy, OmniWeaving can adeptly handle diverse video generation scenarios by weaving free-form text, image, and video inputs into a coherent spatio-temporal narrative.

Meng: The practical implication we see from their focus on reasoning traces is that future AI systems will likely need to incorporate internal planning or tracing mechanisms to handle ambiguous requests effectively. That's a design consideration for us right now.

Lalam: For the future, I think the real impact lies in how this technology allows for more expressive and nuanced content creation, moving past simple instruction following toward true creative partnership with AI systems.

Tom: That's a solid summary of what they're proposing with OmniWeaving: Towards Unified Video Generation with Free-form Composition and Reasoning. We really need to keep an eye on how they refine those reasoning capabilities in the next stages of development.

Conclusion: Tom: So we've heard about OmniWeaving, which is this new framework for making videos that lets you mix text, images, and video in a really free way while having the AI reason through what you want to create.

Jane: Exactly, Tom; it’s about taking the messy parts of video creation—the text prompt and the visual inputs—and stitching them together coherently using this single model structure.

Lu: I think what's really compelling is how they've managed to bake in that reasoning capability directly into the generation process, not just as a separate step afterwards.

Meng: From my side, I'm thinking about how this unified approach could streamline the pipeline; if one model handles everything from the initial thought to the final frame, that cuts down on integration headaches considerably.

Lalam: For me, it’s the potential for creative expression this unlocks; imagine a world where expressing a complex visual idea requires less technical know-how and more just articulating your vision.

Tom: It really boils down to this paper's title, OmniWeaving, suggesting a way to weave these different elements into one seamless fabric.

Jane: The authors are clearly aiming for that unified result by focusing on how the multimodal components talk to each other during the generation process itself.

Lu: They’re addressing the fragmentation issue head-on, showing how to make disparate inputs work together within a single architecture like the MLLM and MMDiT combination.

Meng: I wonder if this level of integration means we'll start seeing more coherent, long-form video content generated from much less structured inputs in the near future.

Lalam: If this holds up, it could fundamentally shift how artists and creators interact with generative systems, making the creative process feel much more intuitive.

Tom: It really suggests that the next generation of video tools won't be single-purpose; they'll be these intelligent agents capable of complex composition on demand.

Jane: The authors’ work points toward a future where video creation becomes less about technical command and more about communicating nuanced intent across multiple media types.

Lu: And this opens up so many avenues for new forms of visual storytelling that we haven't even conceived of yet.

Kaihang Pan, Qi Tian, Jianwei Zhang, Weijie Kong, Jiangfeng Xiong, Yanxin Long, Shixue Zhang, Haiyi Qiu, Tan Wang, Zheqi Lv

Zhejiang University

cs.CV

Submitted: 2026-03-25

Updated: 2026-09-28

Comments: 33 pages, 22 figures. Project Page: https://omniweaving.github.io. Github: https://github.com/Tencent-Hunyuan/OmniWeaving. Model: https://huggingface.co/tencent/HY-OmniWeaving

Code: https://github.com/Tencent-Hunyuan/OmniWeaving

Project page: https://omniweaving.github.io

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 92/100

The gist: OmniWeaving proposes an omni-level video generation model designed to achieve unified, free-form video creation by integrating powerful multimodal composition and reasoning capabilities.

Key concepts

OmniWeaving Architecture
This is the unified framework built on three parts: an MLLM for understanding meaning, a VAE for compressing visual details into low-level signals, and an MMDiT as the main generator. Together, they work to turn mixed inputs into coherent video.
Activating Thinking Mode of the MLLM
Instead of just extracting features passively, this mode makes the MLLM actively reason. It generates intermediate steps to create a more precise prompt before generating video. This enhanced prompt guides the final generation process.
Hidden States DeepStacking
This technique extracts information from many different layers of the MLLM at once. These multi-level features are then fed into the generator to give it richer, multi-granular semantic guidance during video creation.

Terminology

Summary

OmniWeaving proposes an omni-level video generation model designed to achieve unified, free-form video creation by integrating powerful multimodal composition and reasoning capabilities. This model addresses the fragmentation in existing open-source models by learning to temporally bind interleaved text, multi-image, and video inputs while acting as an intelligent agent to infer complex user intentions for sophisticated video creation.

OmniWeaving Architecture

The OmniWeaving framework is built around a unified architecture integrating visual comprehension and generation. It comprises three core components: a Multimodal Large Language Model (MLLM), a Multimodal Diffusion Transformer (MMDiT), and a Variational Autoencoder (VAE). The MLLM acts as the core semantic parser, projecting free-form multimodal inputs into a high-level semantic space. The VAE functions as a visual tokenizer, compressing input visions into low-level latents for fine-grained reconstruction signals. Finally, the MMDiT serves as the backbone diffusion model, where its conditioning branch encodes MLLM semantics and its generative branch integrates VAE latents with noise to generate semantically aligned videos.

Key Model Enhancements

To enhance reasoning and composition, OmniWeaving introduces two specific improvements:

  1. Activating Thinking Mode of the MLLM: This elevates the MLLM from a passive feature extractor to an active reasoner by generating intermediate reasoning steps to autonomously deduce a semantically precise, enhanced prompt. The hidden states of this enhanced prompt are then forwarded alongside the original features to condition the MMDiT.

  2. Hidden States DeepStacking: Inspired by mechanisms in Qwen3-VL, this involves extracting hidden states from a broader range of intermediate MLLM layers to capture a rich semantic spectrum. These multi-level features are projected into the MMDiT embedding space and directly added to the first three layers of the MMDiT conditioning branch, effectively injecting multi-granular semantic guidance into the generative process.

Massive-Scale Training Data

The model is trained on a massive-scale pretraining dataset that encompasses diverse compositional and reasoning-augmented scenarios. This dataset spans both real-world and synthetic domains to capture rich appearances, naturalistic motions, and complex scene dynamics. The training tasks are structured around three primary competencies:

  1. Foundational Video Generation Tasks: These include Text-to-image/video synthesis, Instruction-guided video-to-video editing (for local and global modifications), and Key-frame(s)-to-video generation from a single frame or multiple keyframes.

  2. Multimodal Composition Tasks: These involve Interleaved Text-and-Multi-Image-to-Video generation (where inputs include multiple reference images interleaved with text) and Text-Image-Video-to-Video generation (integrating image, text, and video).

  3. Reasoning-Augmented Tasks: These include Text-to Video generation based on ambiguous queries, Intent-Driven Image-to-Video generation where the model forms a reasoning trace detailing the temporal progression, and Event-Deductive Multi-Image-to-Video generation where it deduces implicit temporal dynamics from disparate keyframes.

IntelligentVBench Benchmark

To rigorously assess unified video generation, OmniWeaving introduces IntelligentVBench, the first comprehensive benchmark designed to rigorously assess next-level intelligent unified video generation. This benchmark employs a VLM-as-a-judge paradigm to evaluate abstract reasoning and compositional capabilities across four distinct tasks: Implicit Image-to-Video (Implicit I2V), Interpolative Dual-Image-to-Video (Interpolative DI2V), Compositional Multi-Image-to-Video (Compositional MI2V), and Text-Image Video-to Video (TIV2V).

Training Strategy

OmniWeaving employs a progressive, three-stage training paradigm:

  1. Stage 1: Modality Alignment Training, where the MLLM parameters are frozen, focusing on fine-tuning the MMDiT with fundamental tasks like Text-to-Video (T2V) and Image-to-Video (I2V).

  2. Stage 2: Multi-Task Free-Form Pretraining, which scales to encompass complex inputs including video editing tasks, partitioning into sub-stages to optimize learning efficiency.

  3. Stage 3: Reasoning-Augmented Fine-Tuning, where the MLLM is unfrozen for end-to-end optimization alongside reasoning tasks, utilizing a next-token-prediction loss specifically designed to enhance the MLLM’s reasoning proficiency.

Evaluation Metrics

The benchmark utilizes three metrics evaluated by Gemini2.5-Pro (VLM-as-a-Judge) on a 1–5 scale for each task:

  1. Instruction Following: Assesses how well the video executes the semantic and logical intent of the text prompt.

  2. Condition Preserving:

Improvements for AI systems

Here are the specific improvements and capabilities derived from the OmniWeaving model and its associated framework, designed to enhance existing video generation systems:


  1. The core improvement is a shift from task-specific modules to an integrated, thinking-guided generalist architecture.

  2. The improved AI system can perform complex, free-form video generation by seamlessly integrating three modalities—text, multiple images, and existing video sequences—into a single coherent spatio-temporal narrative without the need for rigid prompt structuring.

  3. The system excels at Multimodal Composition by intelligently binding disparate inputs:

@ It can take a text prompt alongside several reference images (e.g., defining subjects, scenes, and objects) and synthesize a video where these elements are composed according to complex spatial and temporal relationships (e.g., Place the bun-haired person in a festive sweater at the table, right of the man in red).

  1. The system possesses advanced Abstract Reasoning capabilities:

@ It can act as an intelligent agent to infer complex user intentions from ambiguous or high-level natural language queries (e.g., Two girls were happily greeted by their long-lost dog, or The camera transitions from a traffic light to a historic building). This allows for the generation of videos where the model deduces implicit temporal dynamics and necessary scene changes based on abstract intent rather than explicit, frame-by-frame instructions.

  1. The system enables sophisticated video editing and manipulation:

@ It can perform nuanced operations like Video Editing (e.g., replacing a specific object with one from an image, changing the background entirely, or performing style transfers) while maintaining strict temporal consistency and preserving the motion of all unedited regions. It achieves this by leveraging deep visual understanding to ensure that modifications are physically plausible and seamlessly integrated into the original scene's dynamics.

  1. Training and Evaluation Improvement:

@ The framework introduces a comprehensive evaluation suite, IntelligentVBench, which rigorously tests models on four complex task categories (Implicit I2V, Interpolative DI2V, Compositional MI2V, and TIV2V). This allows researchers to move beyond simple foundational tasks to measure true omni-capable intelligence.

  1. Architectural Enhancement:

@ The model utilizes a DeepStacking mechanism that extracts hidden states from multiple layers of the Multimodal Large Language Model (MLLM) and injects them directly into the generation backbone (MMDiT). This allows for a richer semantic spectrum—spanning from fine-grained details to high-level abstractions—to be used as conditioning, leading to more accurate and higher-fidelity compositional outputs.

  1. Reasoning Enhancement:

@ By activating a Thinking Mode in the MLLM, the system generates intermediate reasoning steps before synthesis. This enables it to autonomously deduce a semantically precise, enhanced prompt that bridges the gap between abstract user intent and pixel-level generation, effectively making the model an active reasoner rather than just a passive renderer.


In summary, OmniWeaving moves video generation from a collection of specialized tools to a single, intelligent agent capable of understanding complex instructions across multiple visual modalities and reasoning through ambiguous scenarios to produce highly coherent, contextually accurate videos.

Sources

Related papers