TwinICL: Diagnosing Multimodal In-Context Learning through Paired Counterfactuals
cs.CV, cs.LG
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: 21 pages, 4 figures
Code: https://github.com/lab-flair/TwinICL
License: http://creativecommons.org/licenses/by/4.0/
The gist: In-context learning (ICL) enables models to infer tasks from demonstrations, but existing benchmarks generally lack matched text and image versions needed to compare ICL performance across modalities.
Terminology
Abstract
In-context learning (ICL) enables models to infer tasks from demonstrations, but existing benchmarks generally lack matched text and image versions needed to compare ICL performance across modalities. We introduce TwinICL, a procedurally generated benchmark providing such pairs for controlled comparison. Across six open-weight models and 38 tasks, multimodal ICL consistently underperforms text-only ICL, with gaps varying by task family. To test whether this gap can be recovered, we target visual access, task framing, and reasoning through three interventions. Their combination recovers strong multimodal ICL performance on a diagnostic subset, despite limited or inconsistent individual effects. To distinguish difficulties in executing tasks from those in inferring them, we evaluate models with explicit task instructions, revealing a modality gap even when the task is known. We then examine how adding demonstration inputs and outputs reshapes this gap, highlighting demonstrations' dual role as additional context to process and evidence about the task. The dataset is available at https://github.com/lab-flair/TwinICL.
Sources
- What do vision-language models see in the context? Investigating multimodal in-context learning
- Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models
- Enhancing Multimodal In-Context Learning via Inductive-Deductive Reasoning
- UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models