MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment
Changhao Xiang, Shangyu Xing, Zhen Wu, Jianbing Zhang, Xinyu Dai
National Key Laboratory for Novel Software Technology, Nanjing University
cs.CV, cs.CL, cs.LG
Submitted: 2026-08-11
Updated: 2026-08-12
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image representations to long textual descriptions.
Terminology
Summary
Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image representations to long textual descriptions. However, this image-level alignment suffers from referential ambiguity: models struggle to infer the correspondences between multiple visual objects and textual entities from the global representation, leading to data inefficiency and suboptimal semantic grounding.
The paper notes that natural images are inherently complex, often containing multiple objects and background elements,
with standard dense caption datasets containing an average of 6.8 to 11.0 distinct entities per sample.
This discrepancy creates a many-to-many
mapping that forces the model to rely on statistical co-occurrences rather than genuine semantic grounding,
making the implicit alignment paradigm highly data-inefficient, necessitating massive-scale datasets to learn robust object-entity associations.
The authors propose MMCS, a novel pretraining paradigm that provides explicit object-level supervision.
Inspired by the linguistic phenomenon of code-switching, MMCS treats vision and language as distinct 'codes'.
Instead of relying solely on global image context, the method create[s] interleaved representations by substituting the embeddings of textual entities with the embeddings of their corresponding visual objects.
Formally, given an original sequence of text tokens X and a contiguous textual entity segment e = [x i,..., x i+m], the interleaved image-text sequence is constructed as:
X MMCS = Concat(Xi+m)
where v object denotes the visual representation of the corresponding object, defined as the set of image tokens whose spatial regions intersect with the object's bounding box.
The training objective combines two losses:
-
L LLM: language modeling loss computed exclusively on text tokens, treating visual tokens as conditioning context
-
L entity: negative log-likelihood of the original textual entity segment given the preceding context and visual object
The overall objective is: L MMCS = L LLM + L entity
The authors develop a scalable data synthesis pipeline to generate a pretraining dataset of 773K samples with accurate object–entity correspondences.
The pipeline consists of:
-
Detailed Image Captioning: Using Qwen3-VL-32B-Instruct to generate detailed captions that
exhaustively describe the attributes of all visible elements within the scene.
-
Textual Entity Extraction: Using Qwen2.5-72B-Instruct to extract noun phrases,
often accompanied by brief attribute descriptions (e.g., a white mug labeled 'O.CO')
for disambiguating multiple objects of the same category. -
Visual Object Grounding: Using Grounding DINO to anchor textual entities to visual regions, outputting bounding box coordinates, text labels, and confidence scores.
-
Filtering: A multi-stage process including box threshold of 0.4 and text threshold of 0.3, pruning bounding boxes that are
smaller than a single visual patch or exceeding 50% of the total image area,
and using SAM-2.1 to discard heavily occluded instances (where mask areas account for less than 20% of bounding-box areas).
The dataset comprises 773,779 samples with 5,145,630 objects, averaging 961.65 characters per caption and 6.65 objects per sample. Quality evaluation shows an average grounding accuracy of 0.949 (VLM judge) and 0.942 (human annotators).
Data Efficiency: With only 50K samples, MMCS achieves downstream performance that surpasses the image-level baseline trained on 600K samples.
This demonstrates extraordinary data efficiency compared to standard image-level pretraining.
Visual Grounding: MMCS achieves an average gain of 7.9%
on RefCOCO/+/g datasets. For example, with Qwen2.5-3B, the average score improves from 59.12 to 70.29.
Visual Perception: MMCS consistently outperforms the baseline with average improvements of 4.0% on CVBench, 2.8% on OCRBench and 2.0% on V-Star.
General VQA: MMCS also improves performance on general VQA tasks, though gains are more modest
because general VQA additionally depends on world knowledge, commonsense reasoning, and instruction-following priors that MMCS does not directly target.
Statistical significance is confirmed with p = 0.028 via paired t-test.
Scaling: MMCS maintains consistent gains when scaling up both dataset size (to 1M samples) and model capacity (Qwen3-8B, Llama3-8B), and generalizes across different vision encoders (SigLIP2 and QwenViT).
MMCS outperforms three baselines:
-
SEA and Patch Aligned Training (patch-level alignment):
aligning semantically complete objects yields cleaner supervision
than patches, whichrarely correspond to complete semantic units, yielding fragmented signals.
-
Text BBox (textualized spatial representation): MMCS
binds visual objects to textual entities at the representation level,
whichproves more effective for modality alignment than symbolic coordinate cues.
-
Removing L entity leads to
the most pronounced drop on perception benchmarks (-4.0%),
as itacts as a semantic anchor
compellingthe projector to encode discriminative, identity-bearing visual features.
-
Removing L LLM causes
the largest drop on grounding benchmarks (-4.9%),
as ittrains the model to predict the textual context surrounding each visual object, encouraging it to reason about how the object relates to its descriptive attributes and to other entities in the scene.
Representation Alignment: Using CKA, CKNNA, and Mutual k-NN metrics, MMCS achieves superior representational alignment compared to image-level supervision
across most layers, supporting the hypothesis that representational alignment correlates with model capability.
Attention Maps: The model pretrained with MMCS accurately attends to visual regions corresponding to specific textual entity descriptions,
while the baseline often produces diffuse or misaligned attention patterns.
Emergent Capabilities: MMCS improves OCR performance despite no dedicated OCR training, attributed to implicit OCR supervision from natural objects
(10.0% to 30.8% of natural objects contain legible text) and enhanced modality alignment.
Learning of Actions and Relationships: Although MMCS only substitutes noun-phrase entities, it also yields notable NLL reductions on verbs and prepositions,
indicating that explicit object-level grounding implicitly propagates to actions and relations.
Token Efficiency: MMCS reduces image token count by approximately 41.4% compared to standard image-level pretraining.
-
Caption quality: MMCS gains persist even when applied to existing ShareGPT-4V captions, confirming improvements
arise from the MMCS formulation itself rather than from the specific teacher pipeline.
-
Noisy correspondences: Under 10% perturbation, MMCS
remains highly robust, still surpassing the caption baseline across all task categories
; even at 30% corruption,the model does not collapse.
-
Image resolution: MMCS
consistently outperforms the baseline across all resolution configurations.
-
Introduction of MMCS,
a novel pretraining paradigm that interleaves visual objects into text, shifting alignment from implicit image-level associations to explicit object-level correspondences.
-
Development of
a scalable data synthesis pipeline that generates 773K samples with precise object-entity correspondences, bypassing the need for manual annotation.
-
Extensive experiments
across various model scales and vision encoders, demonstrating the effectiveness of MMCS,
withmechanistic insights into how MMCS enhances the internal feature space topology of MLLMs.
The current implementation primarily focuses on visual objects within natural images.
Future work includes extending to chart understanding or scene text recognition,
exploring bootstrapping mechanisms for iterative self-improvement,
and extending to encompass complex relations and actions.
Improvements for AI systems
Based on the paper, here are specific improvements I can implement in AI systems:
Improvement: Replace global image-text alignment pretraining with explicit object-entity interleaving (MMCS). During training, substitute textual entity embeddings with corresponding visual object embeddings, using dual losses (L LLM for context prediction + L entity for entity-specific grounding).
What the improved system can do:
-
Resolve referential ambiguity in multi-object scenes (e.g., correctly map
the red mug on the left
to the specific visual region) -
Achieve 7.9% average improvement on referring expression comprehension (RefCOCO/+/g)
-
Learn from 12x fewer samples (50K vs 600K) while surpassing baseline performance
-
Reduce image token usage by 41.4% during pretraining, enabling faster training and lower memory footprint
Improvement: Add L entity loss that forces the visual projector to encode identity-bearing features for each object, not just global scene statistics. This acts as a semantic anchor for discriminative representation learning.
Improvement: Keep L LLM loss on surrounding text tokens, training the model to predict context given visual objects. This teaches relational reasoning between objects and their attributes/actions.
Improvement: Implement the four-stage pipeline: (1) detailed captioning with Qwen3-VL-32B, (2) noun-phrase extraction with Qwen2.5-72B, (3) grounding with Grounding DINO, (4) multi-stage filtering (confidence thresholds, size pruning, occlusion detection via SAM-2.1).
Improvement: Apply MMCS across different model scales (Qwen2.5-3B, Qwen3-8B, Llama3-8B) and vision encoders (SigLIP2, QwenViT) to ensure architecture-agnostic benefits.
Improvement: Design training to tolerate noisy correspondences and variable caption quality, as demonstrated by robustness tests.
Improvement: Leverage implicit supervision from natural objects (10-30.8% contain legible text) to develop OCR capabilities during standard pretraining.
Improvement: Use object-level visual tokens (only those intersecting bounding boxes) instead of full-image tokens during pretraining.
Improvement: Use CKA, CKNNA, and Mutual k-NN metrics to monitor representational alignment between visual and textual modalities during training.
Improvement: Although current implementation focuses on noun-phrase entities, the framework implicitly improves verb and preposition understanding. Extend to explicitly substitute action/relation phrases for deeper semantic grounding.
Abstract
Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image representations to long textual descriptions. However, this image-level alignment suffers from referential ambiguity: models struggle to infer the correspondences between multiple visual objects and textual entities from the global representation, leading to data inefficiency and suboptimal semantic grounding. To address this, we propose MultiModal Code-Switching (MMCS), a novel pretraining paradigm that provides explicit object-level supervision. Inspired by the linguistic phenomenon of code-switching, MMCS interleaves vision and language by replacing textual entities with their corresponding visual objects, enforcing local vision-language grounding. We further develop a scalable data synthesis pipeline to generate a pretraining dataset of 773K samples with accurate object-entity correspondences. Experiments show that MMCS is highly data-efficient: with only 50K samples, it matches or surpasses models trained on 600K image-text pairs. Furthermore, MMCS consistently improves visual grounding and perception capabilities across varying model scales.
Sources
- LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
- Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- Seed1.5-VL Technical Report
- Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models
- DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
- Kosmos-2: Grounding Multimodal Large Language Models to the World
- The Llama 3 Herd of Models
- mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning
- Qwen3 Technical Report
- Qwen2.5 Technical Report
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models