ITO: Multi-View Alignment and Training-Time Fusion for Image-Text Pretraining

arXiv:2603.02767 · cs.CV, cs.AI · Submitted 2026-03-03 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "ITO: Multi-View Alignment and Training-Time Fusion for Image-Text Pretraining".

Tom: ITO (Image–Text as One) is a framework proposed to address modality-induced separation in image-text representations by synergizing multimodal multiple alignment with a lightweight training-time fusion module.

Jane: First, who's behind it and why it matters.

Title and authors: Tom: Moving on to the specifics of who wrote this paper, we're talking about "ITO: Multi-View Alignment and Training-Time Fusion for Image–Text Pretraining." The authors are Hanpeng Liu, Yaqian Li, Zidan Wang, Shuoxi Zhang, Zonglin Zhao, Zihao Bo, Rinyoichi Takezoe, and Kaiwen Long.

Jane: That's quite a team of researchers behind this work. When you see that many authors on a paper like this, it often suggests they’ve been tackling the problem from several different angles to ensure a comprehensive solution.

Lu: I think having multiple authors means they brought diverse expertise to the table, covering both the alignment side and the fusion mechanism design, which is what makes their framework so synergistic.

Meng: So if they’ve covered multiple angles, does that mean their approach is more robust against different types of data noise or biases that might affect one specific part of the system?

Lalam: Absolutely. When you have several perspectives involved, it means they’ve considered potential weaknesses from various viewpoints before settling on the final design, making the resulting framework more resilient.

Tom: Precisely. The paper is structured to tackle this challenge from multiple directions, which is what allows them to build a system that addresses both the alignment and the fusion needs simultaneously.

Jane: So it's not just one person having a single idea, but a collaborative effort where different experts contribute pieces that fit together perfectly.

Lu: Exactly. They built this framework by deliberately combining techniques from different research areas to achieve that unified representation goal, which is what makes the synergy so powerful in ITO: Multi-View Alignment and Training-Time Fusion for Image–Text Pretraining.

Meng: It sounds like they didn't just pick a solution; they engineered it by synthesizing existing ideas into something new that solves the core problem in a novel way.

Lalam: That synthesis is where real innovation happens, and it means we can expect a framework that handles complex multimodal scenarios better than previous iterations.

Tom: So, to recap, the team’s background is diverse enough to handle the complexity of this proposal for ITO: Multi-View Alignment and Training-Time Fusion for Image–Text Pretraining.

The paper's summary: Jane: Now that we know who was behind it, let's get into what they actually propose in terms of the core concepts for "ITO: Multi-View Alignment and Training-Time Fusion for Image–Text Pretraining." What is the main idea they are trying to communicate?

Tom: The main message is that existing image and text contrastive pretraining methods often result in representations that are still partially organized by modality, which ITO aims to fix by introducing two synergistic mechanisms.

Lu: They propose multimodal multiple alignment to enrich supervision by mining diverse image–text correspondences, and then they layer a lightweight training-time multimodal fusion module on top of that.

Meng: So the alignment part is about getting more input diversity, and the fusion part is about making sure those inputs actually talk to each other during optimization.

Lalam: In simple terms, ITO is trying to achieve unified representations by using these two mechanisms together, which means we are aiming for a single space where image and text features share semantic meaning deeply.

Tom: That’s the essence of it. They aren't just improving one aspect; they are restructuring how the model learns the relationship between images and texts from scratch.

Jane: So instead of just adding a layer, they are changing the underlying learning structure itself to enforce that structural interaction during optimization.

Lu: Precisely. They’re moving beyond simple instance-level alignment to one-to-many and many-to-many image–text alignment by exposing the model to diverse cross-modal correspondences during training.

Meng: So it’s about making sure the features learn interactions, not just independent representations that happen to exist in parallel.

Lalam: That structural change is what makes a real difference for our AI because it moves us closer to building systems that possess genuine multimodal understanding rather than just pattern matching.

The paper's improvements: Tom: So we’ve covered the concepts, and now let’s focus on how these two mechanisms actually improve upon what was previously available. What are the specific enhancements they are making to the existing methods?

Jane: Can you explain how their proposed improvements go beyond what simpler alignment-only approaches have achieved in terms of increasing discriminative power?

Lu: The paper states that while multiple alignment improves alignment by enhancing the quality or diversity of individual modality views, it preserves the underlying contrastive supervision structure but doesn't explicitly restructure cross-modal associations within a batch.

Meng: So, what’s the specific limitation that this means in terms of representation organization before we add the second part?

Lalam: That means alignment alone doesn't force those features to interact in a meaningful way across different pairs during training, which is why they needed that fusion module to do something more substantial.

Tom: That’s the exact problem they address: alignment alone isn't enough to fully reshape the organization of the representation space; hence, combining it with training-time fusion acts as a necessary structural regularizer.

Jane: So, what is the crucial improvement offered by that fusion module in terms of stabilizing training dynamics?

Lu: The fusion module pushes apart representations from different pairs while encouraging consistency among representations derived from the same pair, which prevents performance saturation often seen in aggressive contrastive learning. It acts as a regularization against modality-specific overfitting.

Meng: That sounds like it helps keep the model focused on learning general, shared semantics instead of just memorizing specific pairings. It stabilizes the optimization path significantly.

Lalam: For us, that stability is huge because it means our models can scale to larger datasets without hitting that wall where training just stops improving early on due to overfitting.

Tom: And the paper shows this benefit holds even when scaling up to billion-scale datasets; ITO achieves the highest overall performance among all compared methods on DataComp-1B, showing it scales favorably with both training epochs and model size.

Conclusion: Jane: We've covered a lot of ground today, so let’s bring this conversation to a close by summarizing the implications of "ITO: Multi-View Alignment and Training-Time Fusion for Image–Text Pretraining." What are the big implications we should consider for our AI future?

Tom: The core implication is that we can move toward representations that are structurally integrated spaces where image and text features share semantic meaning deeply, which allows for deeper cross-modal reasoning.

Lu: This structural alignment means models are not just aligned but truly integrated at a structural level before they even see their final task instructions.

Meng: This integration suggests that we can build systems that possess genuine multimodal understanding rather than just pattern matching on specific examples.

Lalam: It means our AI is developing features that are more robust and less fragile when faced with novel multimodal inputs or unexpected prompt variations in the future.

Tom: So, to wrap up, "ITO: Multi-View Alignment and Training-Time Fusion for Image–Text Pretraining" shows us how combining multiple alignment with training-time fusion creates a unified representation space that is much more structurally sound.

Jane: It seems like we have a lot of ground to explore as we look ahead at how this impacts the AI landscape, but it’s definitely worth paying attention to this work.

Lu: The potential for creating fundamentally more versatile and integrated models is substantial if they can realize the structural organization they describe.

Meng: From an engineering standpoint, it validates that smart regularization during training can stabilize large-scale pretraining efforts effectively.

Lalam: It gives us a foundation for developing AI systems that are inherently more flexible and less prone to overfitting when scaling up our models to handle the complexity of real-world data.

School of Computer Science and Technology, Huazhong University of Science and Technology

cs.CV, cs.AI

Submitted: 2026-03-03

Updated: 2026-09-30

Code: https://github.com/rom1504/img2dataset2https:

Importance score: 92/100

The gist: ITO (Image–Text as One) is a framework proposed to address modality-induced separation in image-text representations by synergizing multimodal multiple alignment with a lightweight training-time

Key concepts

Multimodal Multiple Alignment
This involves creating diverse image-text pairs from augmented views, such as using two image views and potentially two text views. A contrastive loss is calculated across all these combinations to enrich the supervision signal and improve instance-level alignment beyond conventional setups.
Training-Time Multimodal Fusion Module
This lightweight module processes concatenated visual and textual tokens from an augmented pair using a two-layer Transformer with bidirectional attention. It enforces a 'soft structural constraint' during training, ensuring representations are compatible for deep fusion rather than just linearly separable.
Unified Representation Space
The goal is to reshape the embedding space so that both image and text modalities are mixed within shared neighborhoods, rather than being clearly separated. This structural change makes the resulting embeddings less structured by modality and more organized by shared content.
Inference Efficiency Preservation
The fusion module is only used during training and discarded after training. Consequently, ITO retains the exact architecture of a standard dual-encoder like CLIP, ensuring it has no added parameters, computational cost, or latency during deployment.

Terminology

Summary

ITO (Image–Text as One) is a framework proposed to address modality-induced separation in image-text representations by synergizing multimodal multiple alignment with a lightweight training-time fusion module. This research matters because it demonstrates that while strong alignment drives discriminative power, training-time fusion acts as a critical structural regularizer that unifies the embedding space and stabilizes training dynamics, leading to superior representation quality without sacrificing the efficiency of standard dual-encoder architectures at inference time.

How it works

The framework introduces two synergistic mechanisms: Multimodal multiple alignment and a lightweight training-time multimodal fusion module. The goal is to achieve unified representations through two synergistic mechanisms. The first mechanism, Multimodal multiple alignment, densifies the supervision signal by constructing diverse image–text correspondences from augmented views, which enriches instance-level alignment. Specifically, for each original pair, multiple augmentations are constructed—using two image views and potentially two text views (ITO sub2)—and a bidirectional contrastive loss is computed across all these combinations. This process enriches contrastive supervision beyond conventional setups by exposing the model to multiple image–text correspondences derived from the same underlying sample.

The second mechanism is a lightweight training-time multimodal fusion module designed to enforce structured cross-modal interaction during optimization. Given an augmented pair, visual tokens and textual tokens are concatenated to form a joint sequence, which is then processed by a two-layer Transformer with bidirectional attention to produce fused multimodal tokens. The loss for this fusion objective encourages consistency among representations derived from the same pair while pushing apart representations from different pairs. This soft structural constraint forces encoders to learn features that are not just linearly separable (as in vanilla contrastive learning) but are also compatible for deep fusion, effectively acting as a regularizer against modality-specific overfitting.

Overall Objective and Inference

The final training objective combines the two components: L = LAlign + λLFusion. Here, λ balances the trade-off between discriminative intensity (from alignment) and geometric regularization (from fusion). Crucially, this fusion module is used only during training and discarded at inference, allowing ITO to retain a standard dual-encoder architecture identical to CLIP, thereby preserving the efficiency of standard dual-encoder architectures.

Key Findings on Representation Structure

Analysis reveals a critical interplay between the components: While multiple alignment acts as the primary engine for increasing discriminative power, training-time fusion functions as a necessary structural regularizer. Specifically, ITO consistently eliminates this modality gap and stabilizes training dynamics, preventing performance saturation often observed in aggressive contrastive learning. Furthermore, UMAP visualizations demonstrate that while alignment-only methods like CLIP show a clear separation between image and text embeddings, ITO exhibits a star-shaped distribution in which both modalities are mixed within shared neighborhoods. This confirms that fusion successfully reshapes the organization of the representation space to be less structured by modality and more by shared content.

Impact on Training Dynamics and Scalability

The fusion objective serves as a stabilizer against overfitting. Standard methods like CLIP and SLIP exhibit early saturation (peak accuracy at epoch 26) followed by performance degradation. In contrast, enabling training-time multimodal fusion (λ = 2) stabilizes the training dynamics, leading to consistent performance improvements throughout the full 30-epoch schedule without an early peak. This indicates that fusion is essential for mitigating overfitting and stabilizing contrastive pretraining. Furthermore, experiments on billion-scale datasets show that ITO scales favorably with both training epochs and model size, achieving the highest overall performance among all compared methods on DataComp-1B.

Transferability to Downstream Tasks

Extensive evaluations across various benchmarks confirm the robustness of ITO. Across zero-shot classification, image–text retrieval (on MSCOCO and Flickr30k), and Vision-Language Understanding (following LLaVA-1.5 protocols), ITO consistently outperforms CLIP and strong baselines like SLIP, SigLIP, and FLAIR. The results on complex reasoning tasks suggest that the modality-agnostic structure of ITO’s embedding space significantly lowers the adaptation barrier for Large Language Models, as this improved structural alignment reduces the burden on the projection layer during instruction tuning. Additionally, sub-description sampling (ITO sub2/sub3) further enhances performance when textual supervision is limited, showing that richer and more diverse textual descriptions help our method capture finer semantic correspondences.

Inference Efficiency

The framework maintains high efficiency because the fusion module is discarded at inference time. This means ITO has the same number of parameters, computational cost, and inference latency as CLIP, allowing it to be used as a drop-in replacement for existing image–text contrastive encoders while providing superior representation quality. The training overhead introduced by ITO (SLIP and ITO) is comparable to or slightly higher than CLIP, confirming that the regularization effect is achieved without introducing significant additional computational cost during deployment.

Improvements for AI systems

Based on the scientific paper ITO: Images and Texts as One via Synergizing Multiple Alignment and Training-Time Fusion, here are specific, actionable improvements to AI systems derived from this framework, along with the capabilities these improved systems can achieve:


The core improvement is moving from representations that are merely aligned (modality-separated) to representations that are structurally integrated (unified semantic space). This allows for deeper cross-modal reasoning.

Here are specific improvements and the resulting system capabilities:

  1. Enhanced Cross-Modal Reasoning in Foundation Models:

  2. More Robust Zero-Shot Transfer to Novel Tasks:

  3. Improved Model Stability and Scalability during Training:

The improved AI systems, utilizing the ITO framework, can perform the following specific tasks:

  1. Use a single dual-encoder architecture (like CLIP) for multimodal pretraining but achieve superior performance on downstream tasks because the training process explicitly forces image and text features to be learned in an interleaved, unified space.

  2. Perform complex Visual Question Answering (VQA) or Multimodal Reasoning tasks with higher accuracy, as the model has learned to suppress modality-specific shortcuts and rely on shared semantic content between modalities (e.g., understanding the relationship between a visual scene and its detailed textual description).

  3. Achieve significantly better generalization in zero-shot scenarios (classifying or retrieving) across novel, unseen datasets because the unified representation space is less dependent on modality-specific training signals, leading to more robust feature extraction from new data.

  4. Benefit from improved linear separability of visual features when performing linear probing for classification tasks, meaning the learned visual embeddings are inherently better structured for simple linear classifiers (e.g., distinguishing between different types of objects or scenes).

  5. Maintain high performance and prevent overfitting during large-scale pretraining on massive datasets (like DataComp-1B) by utilizing the training-time fusion module as a structural regularizer, which stabilizes the training dynamics and prevents the early saturation/degradation often seen in aggressive alignment strategies.

  6. Deploy these improved visual encoders in real-world applications (like MLLMs) without increased computational cost or inference latency, as the costly fusion module is discarded at inference time, allowing for direct drop-in replacement of existing dual-encoder backbones.

Sources

Related papers