Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs

arXiv:2604.14520 · cs.CV · Submitted 2026-04-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Chain of Modality".

Jane: Omni-modal Large Language Models (Omni-MLLMs) currently suffer from performance fragility because their static fusion topologies introduce systemic flaws, leading unimodal baselines to frequently outperform joint multimodal inference.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So we're diving into this paper now called "Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs." It sounds like they're tackling a real problem with how current models handle mixing different types of sensory data.

Jane: It definitely has a mouthful of a title, Tom. Essentially, the core idea is moving away from just sticking everything together passively and starting to orchestrate how those different modes work together dynamically depending on what the task is.

Lu: I think it's fascinating because it addresses that performance paradox where unimodal models sometimes do better than multimodal ones even when they have access to multiple inputs. It suggests the fixed ways we fuse things are actually causing problems.

Meng: From my side, I'm curious about how this dynamic orchestration translates into something concrete for deployment; does it just make inference slower, or does it genuinely improve reliability?

Lalam: I see this as a huge step forward for building more resilient AI systems because if the fusion itself is flawed, the whole system is built on sand. This paper suggests fixing that foundation.

The paper's summary: Tom: Okay, so what’s the main takeaway from "Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs"? Essentially, they point out two specific flaws in the current static fusion methods that cause issues.

Jane: They identify "positional bias" when things are fed sequentially, meaning the model relies too much on where something is placed rather than what it actually means, and another issue called an "alignment trap" with interleaved formats where tokens get forced together even if they don't logically belong there.

Lu: That alignment trap sounds like a really subtle problem that can lead to models making mistakes by assuming connections between signals that aren't truly there just because the tokens are adjacent in the input structure.

Meng: So, instead of one fixed way to process inputs, this framework introduces a dynamic system that can switch between parallel, sequential, and interleaved pathways based on what the query actually asks for.

Lalam: That adaptability is key; it moves us from a rigid fusion method to something much more flexible that can choose the best way to look at the data for any given situation.

The paper's improvements: Tom: The framework proposes Chain of Modality, which acts like an agentic system that handles this orchestration. It splits the cognitive workload into a streamlined "Direct-Decide" path for quick tasks and a structured "Reason-Decide" path for more complex analysis.

Jane: That bifurcation is interesting because it means the model can be efficient when it just needs to perceive things intuitively, and then switch to a much deeper analytical mode when it needs to audit evidence.

Lu: The PRD pathway, or Plan-Reason-Decide, where the Reasoner does modality-specific auditing, seems like a very structured way to ensure that any final answer is grounded in the evidence it actually saw.

Meng: I see the efficiency angle here—if the Planner can determine which modalities are truly necessary for a query and prune them out before processing, that should significantly cut down on unnecessary computation.

Lalam: That pruning idea really resonates with me because reducing the computational footprint while maintaining or even improving accuracy is what we need in a practical application; it makes the whole system lighter.

Conclusion: Tom: So, to wrap up, the main point of "Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs" is that dynamic orchestration, rather than static fusion, resolves structural biases like positional bias and alignment traps by switching between parallel, sequential, and interleaved pathways.

Jane: And this leads to the key improvement: bifurcating the model into a streamlined Direct-Decide path for intuitive perception tasks and a more structured Plan-Reason-Decide pathway for complex analytical auditing.

Lu: It suggests that optimizing the physical arrangement of modalities is just as important as designing complex training strategies for unlocking better zero-shot potential.

Meng: I think the practical impact will be seeing models perform reliably across different benchmarks because they aren't relying on a single, potentially flawed fusion method for every input type.

Lalam: This work really shows how we can make the AI more robust by giving it the cognitive tools to dynamically decide how to combine information rather than just passively accepting the input order.

Ziyang Luo, Nian Liu, Junwei Han

Northwestern Polytechnical University

cs.CV

Submitted: 2026-04-16

Updated: 2026-09-28

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 91/100

The gist: Omni-modal Large Language Models (Omni-MLLMs) currently suffer from performance fragility because their static fusion topologies introduce systemic flaws, leading unimodal baselines to frequently

Key concepts

Structural Pathologies in Static Fusion
Static fusion methods suffer from two main problems: positional bias in sequential inputs where visual dominance is an artifact of physical proximity rather than content. Additionally, interleaved formats create an 'alignment trap' by forcing the model to hallucinate semantic consistency between temporally or physically discordant signals.
Chain-of-Modality (CoM) Framework
CoM replaces passive fusion with dynamic orchestration. It treats modality selection as a chain, adaptively switching between parallel, sequential, and interleaved input topologies based on the query's complexity. This ensures the model uses the right interaction style for specific tasks.
Bifurcated Cognitive Execution
CoM splits cognitive work into two paths: a streamlined 'Direct-Decide' path for intuitive tasks that prioritizes perceptual fidelity, and a structured 'Plan-Reason-Decide' pathway for complex analytical auditing. This allows the model to choose the appropriate level of reasoning needed.
Planner Role
The Planner acts as the primary cognitive gatekeeper. It analyzes a query to dynamically determine four key elements: the optimal topological arrangement, cognitive pathway, minimal sufficient modalities, and input format. This planning ensures reasoning is only invoked when necessary.

Terminology

Summary

Omni-modal Large Language Models (Omni-MLLMs) currently suffer from performance fragility because their static fusion topologies introduce systemic flaws, leading unimodal baselines to frequently outperform joint multimodal inference. The Chain of Modality (CoM) framework resolves this by transitioning multimodal fusion from passive concatenation to dynamic orchestration, adaptively switching between parallel, sequential, and interleaved pathways while bifurcating cognitive execution into a streamlined “Direct-Decide” path for direct perception and a structured “Reason-Decide” path for analytical auditing.

Structural Pathologies in Static Fusion

The paper identifies two structural pathologies inherent in current static fusion topologies that cause performance degradation. First, for sequential inputs (e.g., Audio → Visual), the model suffers from positional bias, where visual dominance is an artifact of positional bias, meaning the model relies on structural proximity rather than semantic content. Second, models often adopt interleaved formats to bridge temporal gaps, which introduces an alignment trap, coercing the model into hallucinating semantic consistency between discordant signals due to enforced physical adjacency of cross-modal tokens.

The Chain-of-Modality (CoM) Framework

CoM treats modality selection and interaction as modular building blocks, forming an explicit chain that governs information flow. The framework adaptively orchestrates input topologies based on task complexity, switching among parallel, sequential, and interleaved pathways. This is achieved through a mapping function that determines the optimal topological arrangement based on query intent. For instance, for queries demanding fine-grained temporal alignment (Temporal-centric tasks), CoM constructs an interleaved sequence to capture instantaneous cross-modal correlations. Conversely, for independent evidence auditing (Audio- or Visual-centric queries), it adopts a parallel topology to isolate each modality into separate blocks, preventing the alignment trap.

Bifurcated Cognitive Execution

CoM bifurcates cognitive execution into two task-aligned pathways: a streamlined “Direct-Decide” path for intuitive tasks and a structured “Reason-Decide” path for analytical auditing. For intuitive queries, CoM shortens the chain to a streamlined 'Plan-Decide' (PD) pathway, bypassing generative reasoning to preserve perceptual fidelity. For complex analytical tasks, CoM activates a “Plan-Reason-Decide” (PRD) pathway. In this phase, the Reasoner executes modality-specific evidence auditing, and the Decider synthesizes the resulting logical rationales into a grounded final response by treating rationales as fixed textual tokens to ensure grounding in audited evidence rather than implicit positional biases.

Dynamic Orchestration via Agentic Architecture

The framework reconfigures a single Omni-MLLM backbone into three distinct cognitive roles through specialized system prompting: Planner, Reasoner, and Decider. The Planner serves as the primary cognitive gatekeeper, responsible for decomposing the query to determine the optimal topological arrangement (T), cognitive pathway (P), minimal sufficient modalities (SMmin), and input format (F). This dynamic planning ensures that reasoning is invoked only when cross-modal interactions are necessary, transitioning from a passive aggregator to an active perceiver that dynamically modulates internal attention based on query semantics.

Empirical Validation and Efficiency

Extensive experiments across seven representative benchmarks demonstrate CoM’s superior generalization. The framework achieves consistent, across-the-board gains over vanilla baselines by synergizing a training-free pathway for intuitive tasks with data-efficient SFT for analytical reasoning. Furthermore, the framework exhibits computational efficiency; in scenarios where the Planner identifies redundant modalities, CoM executes modality pruning, effectively reducing the computational footprint compared to the vanilla model. The results confirm that optimizing the physical orchestration of modalities is as vital as complex training designs for unlocking zero-shot potential.

The gist: Chain-of-Modality (CoM) resolves structural pathologies in Omni-MLLMs by dynamically orchestrating modality selection and topology while adaptively routing queries through either a streamlined “Plan-Decide” or a structured “Plan-ReasonDecide” pathway. The paper demonstrates that CoM achieves superior robustness and efficiency across diverse benchmarks by aligning information flow and cognitive depth with task complexity.

How it works

  1. The Planner maps the query to determine the optimal topological arrangement (T), cognitive pathway (P), minimal sufficient modalities (SMmin), and input format (F).

  2. For intuitive tasks, CoM triggers a Direct-Decide flow, mapping organized modalities directly to a response for perceptual fidelity.

  3. For analytical tasks, CoM engages the Plan-Reason-Decide pathway where the Reasoner audits evidence based on the planned topology and generates perceptual rationales.

  4. The Decider synthesizes these rationales into a grounded final response, ensuring logical consensus through decoupled arbitration from raw sensory embeddings.

Improvements for AI systems

Here are the specific improvements to AI systems derived from the Chain-of-Modality (CoM) framework, along with what these improved systems can achieve:


) Implement a dynamic, agentic inference framework called CoM that reconfigures a single Omni-MLLM backbone into three distinct cognitive roles: Planner, Reasoner, and Decider.

  1. The model will first use the Planner to decompose complex queries into minimal sufficient modalities and select an optimal input topology (Parallel, Sequential, or Interleaved) based on task complexity.

  2. For intuitive tasks (e.g., direct perception), CoM activates a streamlined Direct-Decide pathway (Plan-Decide).

  3. For complex analytical queries requiring deep reasoning, CoM activates a structured Reason-Decide pathway (Plan-Reason-Decide).

) This system can overcome the performance paradox where unimodal baselines outperform joint multimodal inference by neutralizing structural biases.

  1. The system will eliminate positional bias in sequential inputs by adaptively choosing input order (e.g., Audio→Visual vs. Visual→Audio) based on task semantics, rather than relying on a fixed default sequence.

  2. The system will resolve alignment traps in interleaved formats by dynamically selecting the topology that best fits the required synchronization level (e.g., using Parallel for independent evidence auditing or Sequential for causal grounding).

) The improved AI system will achieve robust and consistent generalization across diverse benchmarks by synergizing a training-free pathway for intuitive tasks with data-efficient SFT strategies for analytical reasoning.

  1. The system can perform high-fidelity, perception-focused tasks (like identifying instruments from audio clips) with superior accuracy compared to vanilla models, even in zero-shot settings where the Planner dynamically selects the optimal topology.

  2. For challenging analytical and cross-modal reasoning tasks (e.g., verifying semantic consistency between discordant signals), the system can leverage a Plan-Reason-Decide pathway to execute rigorous evidence auditing, leading to significantly higher accuracy on benchmarks like AVHBench and OmniBench, even with limited fine-tuning data.

) The improved AI system will exhibit enhanced computational efficiency through dynamic modality pruning and targeted reasoning activation.

  1. The CoM framework allows the model to prune redundant modalities during the planning stage for specific queries, reducing the computational footprint compared to vanilla models that process the full sensory stream unnecessarily.

  2. By using a fast Direct-Decide pathway for intuitive tasks, the system minimizes unnecessary reasoning noise, ensuring rapid inference time for simple perceptual queries.

Sources

Related papers