MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning

arXiv:2608.12724 · cs.LG · Submitted 2026-08-13 · Read on arXiv

Zirui Cheng, Xun Xu, Tiankai Chen, Fady Rezk, Bowen Zheng, Xiaodong Shi, Shijie Li, Kangkang Lu, Bharadwaj Veeravalli, Nancy F. Chen

National University of Singapore · Institute for Infocomm Research, ASTAR

cs.LG

Submitted: 2026-08-13

Updated: 2026-08-14

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 92/100

The gist: MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning Summary This paper introduces MAG (MAnifold-Guided semi-supervised in-context demonstration selection), a framework designed to

Terminology

Summary

MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning

Summary

This paper introduces MAG (MAnifold-Guided semi-supervised in-context demonstration selection), a framework designed to improve few-shot multi-modal in-context learning (ICL) by leveraging abundant unlabeled data. The authors identify label scarcity as a key bottleneck in multi-modal ICL, where high-quality demonstrations are essential but often limited. They propose a two-stage, training-free approach that uses graph-based relevance score propagation to efficiently select and pseudo-label informative unlabeled samples, then select the most relevant demonstrations for each query.

Problem Setup and Motivation

The paper addresses the challenge that "few-shot in-context learning (ICL) with multi-modal large language models (MLLMs) enables task adaptation without parameter updates, but its performance is highly sensitive to the quality and coverage of the selected demonstrations. While unlabeled multi-modal data is abundant, it remains elusive how to exploit them for ICL. The authors note that in the absence of ground-truth outputs, these unlabeled samples are difficult to exploit using standard retrieval-based ICL methods."

Methodology

MAG formulates demonstration selection as a semi-supervised propagation problem on a multi-modal graph with two stages:

Stage 1: Manifold-Guided Pseudo-Labeling

  • Generates textual descriptions of images using an MLLM: di = Mdesc(xi), where Mdesc produces a detailed description of the visual content

  • Constructs a textual relationship graph over labeled and unlabeled data using text embeddings from Contriever

  • Propagates relevance scores from labeled to unlabeled nodes using the closed-form solution: I(∞) = (1 − α)(I − αŴ(t))−1 I(0)

  • Selects the top-K unlabeled samples with highest relevance scores and pseudo-labels them using the MLLM, reducing inference cost while preserving semantic relevance

Stage 2: Multi-Modal Graph Construction & Query-Conditioned Propagation

  • Constructs separate visual and textual graphs using CLIP ViT-L/14 for visual embeddings and Contriever for text embeddings

  • Propagates query-conditioned relevance scores in both modalities

  • Applies late fusion to combine modality-specific scores: I = βI(v) + (1 − β)I(t)

  • Selects the top-k demonstrations from the expanded candidate pool (labeled + pseudo-labeled)

Key Findings

  1. Textual representations are more effective for Stage 1 relevance propagation: "Textual-only achieves best performance on 6 of 8 datasets, particularly excelling on emotion recognition (EmoSet: 77.0%, Emotion6: 65.0%), scene text understanding (TextOCR: 91.7%), and visual reasoning (MMStar: 70.3%, GQA: 62.3%)."

  2. Both modalities are crucial for Stage 2 demonstration selection: "Removing either modality results in significant performance degradation. Text-only graphs (–img) particularly hurt visually grounded reasoning (GQA: 44.2% vs. 58.0%), while visual-only graphs (–desc) fail on semantically intensive tasks (EmoSet: 60.2% vs. 76.3%, TextOCR: 67.0% vs. 91.4%)."

  3. Late fusion is the most effective fusion strategy: Late fusion consistently achieves the best performance across all benchmarks, with particularly large gains on reasoning-intensive tasks (MMStar: +6.6% over early fusion, GQA: +2.3%).

  4. Manifold-guided pseudo-labeling is critical: Naively introducing pseudo-labels can inject noise and offset potential gains, whereas our relevance-guided selection effectively identifies high-value samples, leading to consistent improvements.

Experimental Results

MAG was evaluated on eight benchmarks spanning visual emotion recognition (EmoSet, Emotion6), scene text understanding (TextOCR), visual reasoning (MMStar, MatchingMI, CLEVR), and visual question answering (GQA, OK-VQA). Using only 15 labeled examples, 585 unlabeled examples, and pseudo-labeling the top 45 samples, MAG consistently outperformed all baselines:

  • Compared to the strongest retrieval baseline (MMICES), MAG achieved +30.5% on CLEVR, +7.1% on MMStar, and +25.1% on TextOCR

  • Relative to standard few-shot prompting, MAG improved TextOCR (+23.8%), CLEVR (+30.8%), and EmoSet (+15.3%)

  • MAG surpassed Top-K+MDL by +34.3% on CLEVR and +27.1% on TextOCR

Ablation Studies

The ablation studies confirmed:

  • Removing pseudo-labeling yields 72.0% on EmoSet, while random pseudo-labels provide marginal gains (73.1%) and can even degrade performance (GQA: 54.7% vs. 56.4%)

  • Graph-based demonstration selection outperforms random selection (+3.5% on EmoSet, +2.5% on TextOCR, +1.8% on GQA)

  • Both modalities are necessary and complementary

Hyperparameter Sensitivity

Performance remains stable across wide parameter ranges with optimal configurations: "α ≈ 0.9 (emphasizing propagation over initialization), knn = 5 (moderate neighborhood size), β = 0.5 (balanced modality fusion), pool size of 60 (15 labeled + 45 pseudo-labeled), label rate of 25% (demonstrating label efficiency), and 10 demonstrations."

Scalability

The method is computationally efficient: "The main computational cost arises from solving the linear system involving (I − αŴ(m)). In practice, Ŵ(m) is a sparse affinity matrix constructed via k-nearest neighbors. Therefore, this reduces to solving a sparse linear system, rather than performing dense matrix inversion." Empirical validation shows near-linear scaling in both latency and memory usage.

Contributions

The paper's contributions are:

  1. Identifying label scarcity as a key bottleneck in few-shot multi-modal ICL and proposing a semi-supervised formulation to leverage unlabeled data

  2. Introducing a two-stage, training-free framework using graph-based relevance score propagation for efficient pseudo-labeling and demonstration selection

  3. Demonstrating consistent gains across 8 diverse benchmarks, with particularly large improvements on reasoning-intensive tasks

The authors conclude that MAG provides a practical foundation for exploiting unlabeled data in future multi-modal ICL systems.

Improvements for AI systems

Improvements to AI Systems:

  1. Semi-Supervised Demonstration Selection for Few-Shot ICL: AI systems can now leverage abundant unlabeled multi-modal data (images + text) without requiring ground-truth labels. By constructing a textual relationship graph and propagating relevance scores from labeled to unlabeled samples, the system automatically identifies and pseudo-labels the most informative examples, expanding the demonstration pool from 15 to 60 samples. This enables task adaptation with as few as 15 labeled examples, dramatically reducing annotation costs.

  2. Training-Free Graph-Based Relevance Propagation: The system can select optimal in-context demonstrations without any fine-tuning or gradient updates. Using closed-form solutions for label propagation on sparse k-NN graphs, it solves a linear system efficiently, achieving near-linear scaling in latency and memory. This allows deployment on resource-constrained devices or in real-time applications where model updates are infeasible.

  3. Multi-Modal Complementary Fusion for Query-Conditioned Retrieval: The AI system now integrates both visual (CLIP ViT-L/14) and textual (Contriever) embeddings via separate graphs, then applies late fusion with a tunable balance parameter (β). This dual-modality approach ensures robust performance across diverse tasks—excelling on visually grounded reasoning (e.g., GQA, CLEVR) and semantically intensive tasks (e.g., emotion recognition, scene text understanding) simultaneously, whereas single-modality systems fail on one or the other.

  4. Noise-Resilient Pseudo-Labeling: The system avoids the pitfall of naive pseudo-label injection (which degrades performance) by using manifold-guided selection to choose only high-value unlabeled samples. This relevance-guided pseudo-labeling consistently improves accuracy (e.g., +3.5% on EmoSet) while reducing inference cost, as only top-K samples (e.g., 45 out of 585) are pseudo-labeled, minimizing computational overhead.

  5. Task-Agnostic and Hyperparameter-Stable Adaptation: The improved system performs reliably across eight diverse benchmarks (emotion recognition, visual reasoning, VQA, scene text) with stable performance over wide hyperparameter ranges (α≈0.9, knn=5, β=0.5, pool size=60). This makes it plug-and-play for new tasks without extensive tuning, enabling rapid deployment in changing environments or novel domains.

  6. Scalable Handling of Large Unlabeled Corpora: The system can process large-scale unlabeled datasets efficiently due to sparse affinity matrices and iterative propagation, avoiding dense matrix inversion. This allows AI systems to continuously ingest new unlabeled data (e.g., user-generated content) and improve demonstration quality over time, supporting lifelong learning scenarios.

What the Improved AI System Can Do:

  • Achieve state-of-the-art few-shot performance on visual emotion recognition (77.0% on EmoSet), scene text understanding (91.7% on TextOCR), and visual reasoning (70.3% on MMStar) using only 15 labeled examples.

  • Outperform existing retrieval-based ICL methods by up to +30.5% on reasoning-heavy tasks (CLEVR) and +25.1% on text-heavy tasks (TextOCR) without any training.

  • Dynamically select the most relevant demonstrations per query by fusing visual and textual relevance scores, adapting to both image-dominant and text-dominant queries in real time.

  • Operate in low-label regimes (e.g., 25% label rate) with minimal performance loss, making it viable for niche or emerging domains where labeled data is scarce.

  • Provide a practical, training-free foundation for exploiting unlabeled multi-modal data in production systems, reducing the need for human annotation and enabling faster iteration on new tasks.

Sources

Related papers