FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens

arXiv:2506.03096 · cs.CV, cs.LG · Submitted 2026-08-15 · Read on arXiv

Christian Schlarmann, Francesco Croce, Nicolas Flammarion, Matthias Hein

University of Tübingen · EPFL

cs.CV, cs.LG

Submitted: 2026-08-15

Updated: 2026-08-18

Comments: Code and models available at https://github.com/chs20/fuselip

Code: https://github.com/chs20/fuselip

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 76/100

The gist: FuseLIP introduces a novel architecture for multimodal embedding that extends CLIP to handle multimodal inputs, such as image-text pairs, by encoding them into a single feature vector.

Terminology

Summary

FuseLIP introduces a novel architecture for multimodal embedding that extends CLIP to handle multimodal inputs, such as image-text pairs, by encoding them into a single feature vector. The key innovation is early fusion: instead of using separate encoders for each modality and merging their features late, FuseLIP uses a single transformer encoder that operates on a unified sequence of discrete tokens from both text and image. This is achieved by leveraging a frozen discrete image tokenizer from the TiTok family, which, together with a standard text tokenizer, maps inputs into tokens from a finite multimodal vocabulary. The paper states: "we propose to use a single transformer model which operates on an extended vocabulary of text and image tokens. This early fusion approach allows the different modalities to interact at each depth of encoding and obtain richer representations compared to common late fusion." The architecture processes tokenized inputs with a single encoder, and the final embedding is derived from the output of a special end-of-text token. This design allows the model to be trained with a contrastive loss similar to standard CLIP, despite using a single encoder.

The training objective combines a sigmoid contrastive loss, adapted from SigLIP, with a masked multimodal modeling (MMM) loss. The MMM loss is applied by randomly masking tokens and training the model to predict them, which is made straightforward by the discrete tokenization. The paper notes: the masked modeling and contrastive losses are applied on the same masked input, avoiding extra computational overhead. The final loss is a weighted combination of the two, with the MMM loss weight set to 0.25. The image tokenizer remains frozen during training.

To train and evaluate the model, the authors collect new datasets. For training, they use unimodal image-text pairs from CC3M and CC12M, and generate multimodal data. This includes text-guided image transformations (TGIT), where images are transformed (e.g., cropped, rotated, flipped) and paired with textual descriptions of the transformation. They also generate VQA data from captions using an LLM, and use existing VQA data from Visual Genome. Additionally, they create a visual grounding dataset (VG-Crop) from Visual Genome region descriptions and use the HQ-Edit dataset for image editing tasks. The paper emphasizes the importance of hard negatives: We design and integrate hard negatives into training of FuseLIP and baselines by ensuring that batches contain semantically similar examples. This is done by sampling multiple transformations of the same image, or multiple samples from the same query image.

For evaluation, the authors use existing benchmarks like MMEB and ImageNet, and introduce new tasks. These include OI-Crop and OI-Pos, created from OpenImages, which test the ability to select the correct crop of an object given a text query, and to distinguish between left and right instances of an object. They also evaluate on VG-Crop and CC3M-TGIT. The paper reports that FuseLIP-B achieves the best results across nearly all tasks, often with a large margin, compared to late fusion baselines (score fusion and MagicLens-style fusion). The largest improvements are seen on CC3M-TGIT, where FuseLIP outperforms baselines by 9-24%, particularly on tasks requiring identification of the correct image after cropping, rotation, or flipping. The paper explains: "The advantage of FuseLIP emerges specifically in tasks requiring identification of the correct image after cropping, rotation, or flipping... This improvement likely stems from the nature of these tasks, which rely on capturing the visual structure rather than semantic content."

Ablation studies show that both hard negatives and the MMM loss are crucial for performance. Removing hard negatives causes large drops on VG-Crop, OI-Crop, and CC3M-TGIT, while removing the MMM loss leads to significantly worse results across all tasks. The paper concludes that early fusion of discrete tokens is highly effective for multimodal representation learning, and that the proposed architecture simplifies the combination of contrastive and masked modeling objectives. The authors also note limitations, including limited computational resources preventing tests at larger scale, and higher inference cost compared to baselines.

Improvements for AI systems

Based on the paper, here are specific improvements I can implement in AI systems:

Improvement: Replace separate text/image encoders with a single transformer that processes discrete tokens from both modalities (images tokenized via TiTok, text via standard tokenizer) in a unified vocabulary.

Capability: The system can now:

  • Encode image-text pairs into a single feature vector without additional fusion modules

  • Allow modalities to interact at every transformer layer (early fusion)

  • Handle unimodal inputs (text-only, image-only) and multimodal inputs (image+text) with the same model

Improvement: Train with both SigLIP contrastive loss and masked token prediction (MMM loss) on the same forward pass, with α=0.25 weighting.

Improvement: Construct batches containing semantically similar samples (multiple transformations of same image, multiple crops from same image, inverse edits) to serve as hard negatives.

Improvement: Train on generated datasets (CC3M-TGIT, CC12M-TGIT) with transformation descriptions paired with transformed images.

Improvement: Use frozen TiTok tokenizer (trained only for reconstruction, not text-alignment) to avoid bias and reduce trainable parameters.

Improvement: The early fusion architecture naturally captures compositional relationships better than late fusion.

Improvement: The model can directly retrieve images based on image+text queries without fine-tuning.

Improvement: Generate training data from existing image-caption datasets using LLMs (for VQA) and automated transformations (for TGIT).

These improvements enable AI systems to handle complex multimodal reasoning tasks that require understanding both visual details and textual instructions simultaneously, which is crucial for applications like visual question answering, image editing, robotics navigation, and assistive technologies.

Abstract

Contrastive language-image pre-training aligns the features of text-image pairs in a common latent space via distinct encoders for each modality. While this approach achieves impressive performance in several zero-shot tasks, it cannot natively handle multimodal inputs, i.e., encoding image and text into a single feature vector. As a remedy, it is common practice to use additional modules to merge the features extracted by the unimodal encoders. In this work, we present FuseLIP, an alternative architecture for multimodal embedding. Leveraging recent progress in discrete image tokenizers, we propose to use a single transformer model which operates on an extended vocabulary of text and image tokens. This early fusion approach allows the different modalities to interact at each depth of encoding and obtain richer representations compared to common late fusion. We collect new datasets for multimodal pre-training and evaluation, designing challenging tasks for multimodal encoder models. We show that FuseLIP outperforms other approaches in multimodal embedding tasks such as VQA and text-guided image transformation retrieval, while being comparable to baselines on unimodal tasks.

Sources

Related papers